Architecture

Thin client. Hosted inference.

The whole point is that you should not need local compute to use N-ATLaS. All inference runs on a hosted GPU endpoint. What you install is a lightweight TypeScript client.

Your application

Citizen services, education, customer service, or anything custom.

new OpenAtlas({ apiKey })

OpenAtlas SDK npm

One client class, one config object. Routes chat() to the LLM, transcribe() to the right one of four ASR models, and speak() to TTS (stretch). Repairs text locally with normalizeText() and sends corrections with reportIssue(). Handles retries, timeouts and typed errors.

OpenAtlas gateway Cloudflare Worker

Checks your key, counts active end users against the licence cap, stores corrections from reportIssue(), and forwards each call to the GPU host. The host's credentials never leave it.

GPU host · AMD Instinct MI300X (live demo), or a free Kaggle notebook, or any GPU you run it on

N-ATLaS LLM

Llama-3 8B fine-tune, served in bf16 on the MI300X (4-bit on 16 GB GPUs such as Kaggle's T4).

N-ATLaS ASR ×4

Hausa, Yoruba, Igbo and Nigerian-accented English.

TTS Optional

SoroTTS for single sentences, MMS-TTS for longer text and English. Final audio rendering only; reliable in English.

Why this shape

Two deliberate choices.

A thin client

A local-first design would contradict the core value proposition: no GPU required to use N-ATLaS. So inference is centralized and the client is kept deliberately small.

A gateway in front of a swappable host

GPU hosts give each endpoint its own URL and key, and count requests rather than people. The gateway gives developers one URL and their own key, measures unique end users against the licence cap, and lets the host move with a configuration change. It was built and proven on a free Kaggle notebook first; the live demo runs the same server on an AMD Instinct MI300X (DigitalOcean AMD Developer Cloud), started with one command. Anyone can stand up the same backend with the bootstrap script, the Kaggle notebook or the Docker image (see Deploy your own in the docs).

Components

What does what.

ComponentTechnologyResponsibility
Inference hostAMD Instinct MI300X on DigitalOcean (ROCm, live demo); a free Kaggle T4 x2 notebook; or Docker on any NVIDIA GPURuns the N-ATLaS LLM and four ASR models and exposes HTTP endpoints.
Model servingOne Python server (FastAPI and transformers) for the LLM and all four ASR modelsTurns handler calls into model forward passes.
TTS hostSame deployment, separate endpointRuns SoroTTS or MMS-TTS for audio rendering. Optional; reliable in English only.
OpenAtlas SDKTypeScript, published to npmOne client over all of the above, written fresh with the inference settings tested against the real N-ATLaS weights.
GatewayCloudflare Worker and D1One public URL and per-developer keys. Counts active end users against the licence cap, stores reportIssue() corrections and holds the GPU host's credentials.
Starter kits ×3Node and TypeScriptMinimal reference apps proving the SDK end to end.
DocsMarkdown README and TSDocQuickstart and API reference.

Infrastructure and compliance

How it's deployed and kept honest.

Compute

One GPU process serves the LLM and all four ASR models, which fit together on a 16 GB GPU. An optional separate service for TTS.

Model storage

Weights are pulled from Hugging Face when the backend starts.

Distribution

An npm package named openatlas, or the nearest available name, with source on GitHub.

Secrets

Per-developer OpenAtlas keys, stored only as hashes. The GPU host's credentials stay in the gateway. Per-key quotas are a roadmap item.

Licence compliance

Positioned as a non-commercial resource. The gateway counts distinct end users over a rolling 30 days and refuses new ones at the 1,000 cap.

Data and transport

All SDK traffic over HTTPS. Starter kits handle demo-scale, ephemeral data only, with no persistence layer.

Build sequencing

Ten days, in parallel.

The plan to the October 12 submission. A live N-ATLaS backend is the single hard dependency, so it comes first. Speech output is gated behind an explicit go or no-go by October 7.

DaysWork
Oct 2Docs repositioned around the SDK. Backend server and hosting notebook. Gateway made host-agnostic. normalizeText() and reportIssue().
Oct 2–3N-ATLaS on a free notebook GPU (Kaggle). First real chat() and transcribe() calls through the gateway.
Oct 4Backend moved to an AMD Instinct MI300X (ROCm): full precision, one-command bootstrap, every check re-run.
Oct 3–4Move the backend to NiHub with a configuration change, then re-verify everything.
Oct 4–7Starter kits and website against the real SDK. Check for the NAIC reply; go or no-go on TTS.
Oct 8–9If confirmed, build speak(). If not, harden the core. End-to-end test of every kit on the persistent host, with no mocks.
Oct 10–11Publish the SDK, write the submission and demo script. A buffer, not a scope-adding one.
Oct 12Submit.

After submission

What comes next

Reasonable engineering goals, explicitly not demo-day claims.

01

Python SDK

A port beyond TypeScript and Node.

02

Per-key quotas

Add per-developer rate limits on top of per-developer keys.

03

Load testing

Measure reliability and multi-tenant behavior before any uptime claim.

04

Cost efficiency

Tune GPU tiers and cold-start behavior against real traffic.

05

Tone restoration

Use N-ATLaS itself to restore tone marks that were never typed. Roadmap only.