Compute
One GPU process serves the LLM and all four ASR models, which fit together on a 16 GB GPU. An optional separate service for TTS.
Architecture
The whole point is that you should not need local compute to use N-ATLaS. All inference runs on a hosted GPU endpoint. What you install is a lightweight TypeScript client.
Citizen services, education, customer service, or anything custom.
new OpenAtlas({ apiKey })
One client class, one config object. Routes chat() to the LLM, transcribe() to the right one of four ASR models, and speak() to TTS (stretch). Repairs text locally with normalizeText() and sends corrections with reportIssue(). Handles retries, timeouts and typed errors.
Checks your key, counts active end users against the licence cap, stores corrections from reportIssue(), and forwards each call to the GPU host. The host's credentials never leave it.
Llama-3 8B fine-tune, served in bf16 on the MI300X (4-bit on 16 GB GPUs such as Kaggle's T4).
Hausa, Yoruba, Igbo and Nigerian-accented English.
SoroTTS for single sentences, MMS-TTS for longer text and English. Final audio rendering only; reliable in English.
Why this shape
A local-first design would contradict the core value proposition: no GPU required to use N-ATLaS. So inference is centralized and the client is kept deliberately small.
GPU hosts give each endpoint its own URL and key, and count requests rather than people. The gateway gives developers one URL and their own key, measures unique end users against the licence cap, and lets the host move with a configuration change. It was built and proven on a free Kaggle notebook first; the live demo runs the same server on an AMD Instinct MI300X (DigitalOcean AMD Developer Cloud), started with one command. Anyone can stand up the same backend with the bootstrap script, the Kaggle notebook or the Docker image (see Deploy your own in the docs).
Components
| Component | Technology | Responsibility |
|---|---|---|
| Inference host | AMD Instinct MI300X on DigitalOcean (ROCm, live demo); a free Kaggle T4 x2 notebook; or Docker on any NVIDIA GPU | Runs the N-ATLaS LLM and four ASR models and exposes HTTP endpoints. |
| Model serving | One Python server (FastAPI and transformers) for the LLM and all four ASR models | Turns handler calls into model forward passes. |
| TTS host | Same deployment, separate endpoint | Runs SoroTTS or MMS-TTS for audio rendering. Optional; reliable in English only. |
| OpenAtlas SDK | TypeScript, published to npm | One client over all of the above, written fresh with the inference settings tested against the real N-ATLaS weights. |
| Gateway | Cloudflare Worker and D1 | One public URL and per-developer keys. Counts active end users against the licence cap, stores reportIssue() corrections and holds the GPU host's credentials. |
| Starter kits ×3 | Node and TypeScript | Minimal reference apps proving the SDK end to end. |
| Docs | Markdown README and TSDoc | Quickstart and API reference. |
Infrastructure and compliance
One GPU process serves the LLM and all four ASR models, which fit together on a 16 GB GPU. An optional separate service for TTS.
Weights are pulled from Hugging Face when the backend starts.
An npm package named openatlas, or the nearest available name, with source on GitHub.
Per-developer OpenAtlas keys, stored only as hashes. The GPU host's credentials stay in the gateway. Per-key quotas are a roadmap item.
Positioned as a non-commercial resource. The gateway counts distinct end users over a rolling 30 days and refuses new ones at the 1,000 cap.
All SDK traffic over HTTPS. Starter kits handle demo-scale, ephemeral data only, with no persistence layer.
Build sequencing
The plan to the October 12 submission. A live N-ATLaS backend is the single hard dependency, so it comes first. Speech output is gated behind an explicit go or no-go by October 7.
| Days | Work |
|---|---|
| Oct 2 | Docs repositioned around the SDK. Backend server and hosting notebook. Gateway made host-agnostic. normalizeText() and reportIssue(). |
| Oct 2–3 | N-ATLaS on a free notebook GPU (Kaggle). First real chat() and transcribe() calls through the gateway. |
| Oct 4 | Backend moved to an AMD Instinct MI300X (ROCm): full precision, one-command bootstrap, every check re-run. |
| Oct 3–4 | Move the backend to NiHub with a configuration change, then re-verify everything. |
| Oct 4–7 | Starter kits and website against the real SDK. Check for the NAIC reply; go or no-go on TTS. |
| Oct 8–9 | If confirmed, build speak(). If not, harden the core. End-to-end test of every kit on the persistent host, with no mocks. |
| Oct 10–11 | Publish the SDK, write the submission and demo script. A buffer, not a scope-adding one. |
| Oct 12 | Submit. |
After submission
Reasonable engineering goals, explicitly not demo-day claims.
A port beyond TypeScript and Node.
Add per-developer rate limits on top of per-developer keys.
Measure reliability and multi-tenant behavior before any uptime claim.
Tune GPU tiers and cold-start behavior against real traffic.
Use N-ATLaS itself to restore tone marks that were never typed. Roadmap only.