Documentation
Install it. Call it. Build on it.
Everything you need to get a real response from a real N-ATLaS model: quickstart, API reference, language codes and the limits to know about.
Quickstart
From zero to a response.
Three steps. No GPU, no model weights, no local setup beyond Node.
Install the SDK
# from your project
npm install @openatlas/sdkRequest an API key
Keys are issued on request. There is no dashboard yet. Request one through the key request form, then export it.
export OPENATLAS_API_KEY="your-key"Make your first call
Ask a question in Yoruba. Copy, paste, run.
import { OpenAtlas } from "@openatlas/sdk"; const client = new OpenAtlas({ apiKey: process.env.OPENATLAS_API_KEY }); const response = await client.chat({ messages: [{ role: "user", content: "Ṣe o le ṣàlàyé ìdí tí ọ̀run fi jẹ́ búlúù?" }], user: "your-end-user-id", // required: one ID per end user }); console.log(response.content);
First call slow?
The first request after the backend starts waits while the models load. See What to expect below.
API reference
Five methods.
Request and response shapes deliberately echo familiar chat-completion conventions. Every model call except the optional speech endpoint is served by an N-ATLaS model.
Text reasoning via the N-ATLaS LLM. Pass a messages array; language is an optional hint.
| Field | Type | Notes |
|---|---|---|
| messages | { role, content }[] | Required. |
| user | string | Required. A stable, opaque ID for your end user. Hashed, and used only to count active users against the licence cap. |
| language | "ha" | "yo" | "ig" | "en" | Optional. |
| returns | { content, model, attribution, usage } | model is "NCAIR1/N-ATLaS"; attribution is "Powered by Awarri". |
Speech to text. The SDK routes to the matching N-ATLaS ASR model by language code.
const result = await client.transcribe({ audio: audioBase64, language: "yo", user: currentUser.id, }); console.log(result.text);
| Field | Type | Notes |
|---|---|---|
| audio | bytes or base64 | Required. Any common format (wav, mp3, ogg, webm, m4a), about 7 MB at most. |
| language | "ha" | "yo" | "ig" | "en-ng" | Required. Selects the ASR model. |
| user | string | Required. Same as chat(). |
| returns | { text, language, model, attribution } | model is the ASR model's ID, e.g. "NCAIR1/Hausa-ASR"; attribution is "Powered by Awarri". |
Repairs Nigerian-language text whose special characters were corrupted by typing or scraping: encoding damage (Æ™asa to ƙasa), look-alike letters (Ş to Ṣ for Yoruba, ķ to ƙ for Hausa), invisible characters and Unicode composition. Hausa apostrophe spellings such as k'asa to ƙasa are opt-in. Also available as client.normalizeText().
import { normalizeText } from "@openatlas/sdk"; normalizeText("Ina zan sabunta katin Æ™asa?", { language: "ha" }); // → "Ina zan sabunta katin ƙasa?"
Repair, not restoration
It does not add tone marks that were never typed: missing marks stay missing. N-ATLaS-based tone restoration is a roadmap item.
Flags a wrong N-ATLaS output with its correction. Reports are stored by OpenAtlas as an exportable correction dataset. Nothing is stored unless you call it, so tell your users when you do.
| Field | Type | Notes |
|---|---|---|
| kind | "chat" | "transcription" | Required. |
| output | string | Required. What N-ATLaS returned. |
| correction | string | Required. What it should have been. |
| input | string | Required for chat: the prompt. |
| audio | bytes or base64 | Transcription only, up to about 1 MB encoded, so the corrected transcript is paired with its audio. |
| returns | { id, received_at } |
Final-stage audio rendering of text N-ATLaS already produced, by a separate text-to-speech model: SoroTTS for single sentences, Meta MMS-TTS for longer text and English. It is on for the hosted API and has been verified on the real models. English comes back clearly: 2–5% of words wrong when the audio is played back through the N-ATLaS speech recognizer. Hausa, Yoruba and Igbo are rendered, but with 36–79% wrong, so treat them as experimental. A gateway with it switched off returns tts_disabled. The website's own demos don't use it.
| Field | Type | Notes |
|---|---|---|
| text | string | Required, up to 1,000 characters. Spoken unchanged. |
| language | "en" | "ha" | "yo" | "ig" | "pcm" | Required. "pcm" is Nigerian Pidgin. |
| engine | "auto" | "sorotts" | "mms" | Default "auto": SoroTTS for a single sentence where it covers the language (natural, slow); MMS-TTS (fast) for longer text, English, or on failure. |
| user | string | Required, as for chat(). |
| returns | { audio, seconds, engine, model, attribution, warnings } | audio is WAV bytes. model is the TTS model's Hugging Face ID, e.g. "Shinzmann/sorotts" or "facebook/mms-tts-hau". attribution is its license credit. |
A strict boundary
N-ATLaS has no text-to-speech model. speak() only renders audio. It never reasons, transcribes or replaces N-ATLaS at any decision point.
Language codes
Which code calls which model.
| Code | Language | transcribe() | chat() |
|---|---|---|---|
| ha | Hausa | N-ATLaS ASR, Hausa | Accepted |
| yo | Yoruba | N-ATLaS ASR, Yoruba | Accepted |
| ig | Igbo | N-ATLaS ASR, Igbo | Accepted |
| en-ng | Nigerian-accented English | N-ATLaS ASR, English | Use "en" |
Errors
Typed and readable.
Errors are named and human-readable, not raw HTTP stack traces, so you know what happened and what to do next.
| Error | Meaning | What to do |
|---|---|---|
| OpenAtlasAPIError | The endpoint returned an error response. | Check the message and your request. |
| OpenAtlasTimeoutError | The request timed out, often while models are loading. | Retry, and expect first-call latency. |
| OpenAtlasConnectionError | The gateway could not be reached. | Check your network and the base URL. |
| OpenAtlasError | Invalid input caught before any request, such as a missing user. | Fix the call; nothing was sent. |
What to expect
Cold starts are normal.
When the backend has just started, the first call waits while the models load. Later calls are faster. This is documented here rather than discovered the hard way.
Known limitations
Stated plainly.
| Licence | N-ATLaS is non-commercial and capped at 1,000 active users per 30 days. Basic request logging is kept to monitor usage against that cap. |
| Reliability | New infrastructure, not load-tested, no uptime or SLA claim. |
| Speech output | speak() is built but switched off on the hosted gateway until it has been verified against the real speech models. It never changes the text it is given. |
| Text repair | normalizeText() fixes corrupted characters. It does not restore tone marks that were never typed; N-ATLaS-based restoration is a roadmap item. |
| Corrections | reportIssue() collects and exports corrections. There is no agreed hand-off to the N-ATLaS maintainers yet. |
| Languages of the SDK | TypeScript and Node only. A Python SDK is a roadmap item. |
Deploy your own
Run the whole stack yourself.
Two halves:
- Backend: a GPU host running the N-ATLaS LLM and the four speech-recognition models.
- Gateway: a Cloudflare Worker with a D1 database, which holds the keys, counts users against the license cap and forwards requests.
We built and proved the whole stack on free tools first: a Kaggle notebook for the GPU, a Cloudflare quick tunnel, and the Cloudflare free plan for the gateway. The live demo runs the same server on an AMD Instinct MI300X (DigitalOcean AMD Developer Cloud). Anyone can stand up the same working deployment with the steps below. Moving the backend is a config change on the gateway, not a code change. The full guide is docs/deploy-your-own.md in the repository.
| Path | Status |
|---|---|
| AMD Developer Cloud (MI300X) | Proven; the live host. One command (deploy/amd/up.mjs) from a freshly created droplet to serving in under 5 minutes, models included. |
| Kaggle notebook backend + gateway | Proven; the free path (T4 x2). |
| One-command gateway setup | Checked with a dry run only; not yet run against a fresh Cloudflare account. |
| Docker on your own GPU | The image builds and starts in CI (no GPU there); not yet run on a GPU by us. |
| Colab notebook | Earlier proven host (Oct 2–3); same code, but not re-run since the latest notebook changes. |
| RunPod Serverless | Scripted, never run. Its ASR worker still uses the older audio chunking, so keep requests to 30 s or less. |
Get access
- Hugging Face: accept the terms on NCAIR1/N-ATLaS and the four NCAIR1 ASR repos, then create a read token.
- Cloudflare: an API token on the free plan with Workers Scripts Edit and D1 Edit.
Start the backend on Kaggle (free)
- Import deploy/colab/natlas_kaggle.ipynb into Kaggle. Under Settings, choose GPU T4 x2 and turn Internet on (this needs a phone-verified account).
- Under Add-ons → Secrets, add HF_TOKEN and NATLAS_API_KEY (any long random string), and attach both.
- Run all. It loads and tests every model, then prints a public URL. Leave the tab open.
Have your own GPU? The Docker image runs the same server; see the guide.
Set up the gateway in one command
CLOUDFLARE_API_TOKEN=… node deploy/setup-gateway.mjs --name my-openatlas \ --backend <NATLAS_BASE_URL> <NATLAS_API_KEY> # creates the database, deploys, sets an admin token, connects the backend, # issues your first key and writes .env.selfhost. Add --dry-run to preview.
Check it with a real answer
node --env-file=.env.selfhost sdk/examples/quickstart.mjs
Your deployment, your obligations
N-ATLaS is non-commercial and capped at 1,000 active users per 30 days; your gateway enforces the cap. Show "Powered by Awarri" with model output. A free Kaggle session ends after at most 12 hours (and uses the account's weekly GPU quota), and its URL changes each run. Reconnect with deploy/set-backend.mjs; no redeploy is needed.