Runtime Documentation
Everything the sofuu binary exposes: the CLI commands, the interactive agent, and the JavaScript modules available in every script.
Installation
One command installs Sofuu on macOS (Apple Silicon), Linux (x86_64 and arm64), or Windows (x64) — a single static executable with zero dependencies, shipping with QTSQ encrypted local persistence. The one-line installer puts sofuu on your PATH so it works from any terminal in any project.
# Install globally onto your PATH — macOS, Linux, or Windows$ curl -fsSL https://sofuu.xyz/install | sh# Then, from any directory:$ sofuu run app.ts# or via npm — the same binary, verified against its published sha256$ npx sofuu
Quickstart Guide
Sofuu runs .js files on an embedded QuickJS engine — instant boot, no JIT warmup, no install step. Here is a complete edge server that streams a local LLM to the browser:
// app.js — an edge server that streams a local LLM to the browserconst server = sofuu.http.createServer((req, res) => {if (req.url === "/chat") {const stream = sofuu.ai.stream("Explain quantum mechanics.", {provider: "ollama", model: "llama3"});res.writeHead(200, { "Content-Type": "text/event-stream" });// chunks arrive from the native polling layer — the loop never blocksfor await (const chunk of stream) res.write(chunk.text);res.end();} else {res.writeHead(200);res.end("AI Edge Runtime Active");}});server.listen(8080, "0.0.0.0");console.log("Server bound on port 8080.");
Run it from your terminal: $ sofuu run app.js. Point provider at any local or cloud LLM — configuration lives in /provider inside the chat, or pass it per call as shown.
CLI Commands
One binary is the whole toolchain: runtime, package manager, agent host, and REPL.
| Command | What it does |
|---|---|
sofuu | Boots the interactive Read-Eval-Print-Loop (REPL). |
sofuu run <file.js|ts> | Executes a script via the embedded QuickJS engine — no build step, no warmup. |
sofuu eval "code" | Evaluates a JS string inline from any shell. |
sofuu chat | Starts the interactive AI agent — see the next section. |
sofuu bundle <entry> -o <out> | Bundles your files and modules into a single distributable script. |
sofuu install | Resolves package.json dependencies without deep node_modules trees. |
sofuu add <pkg> | Installs a module from the registry. |
sofuu doctor | Health check: confirms the local brain store is linked and round-trips a write, shows where it lives, and prints the resolved context window for the active model with the evidence behind it. |
sofuu version | Prints the runtime version and platform. |
Chat & the Agent
sofuu chat is a full coding agent in your terminal: streaming answers and a built-in tool belt — read, write, edit, grep, glob, list_dir, and bash — with sessions that persist locally as encrypted .qtsq containers. Nothing leaves your machine except the model calls themselves.
# Start the interactive agent$ sofuu chat# Pick a provider and model on first run, then just talk to it.# It can read and edit files, search code, and it remembers# across sessions in an encrypted local store.
The most useful slash commands:
| Command | Purpose |
|---|---|
/provider · /model | Pick the LLM backend and model — local (Ollama) or cloud. |
/apikey · /baseurl | Store credentials for a provider, outside your scripts. |
/effort | Set the model's reasoning effort. |
/tools | Toggle the built-in tool belt on or off. |
/ctx · /cost | Live context-window and spend meters for the session. Accepts 1m, 128k or a raw token count; an explicit value is honored as typed, and /model detects the window from the endpoint automatically. |
/compact | Compact the conversation when the context window fills up. |
/remember | Save a durable memory to the local brain store. |
/sessions · /resume | List and continue past sessions. |
/help | The full command list. |
Global Sandbox
Primitives injected into the root scope of every script — no import needed.
| Binding | Behavior |
|---|---|
setTimeout(cb, ms) | Schedules the callback on the native event loop. Cancel with clearTimeout(id). |
setInterval(cb, ms) | Repeating background scheduler. Cancel with clearInterval(id). |
sleep(ms) | Synchronous halt of the entire event loop — useful in embedded scripts, avoid in servers. |
console.log() | Standard output; console.time(id) / timeEnd(id) for high-resolution timing. |
process.argv · process.env | Command-line arguments and the process environment block. |
process.cwd() | Current working directory. |
process.exit(code) | Terminates the runtime immediately. |
process.stdout.write(s) | Raw write to standard output. |
sofuu.ai
The AI bridge: talk to local models (Ollama) or cloud providers (OpenAI, Anthropic, OpenRouter, …) through one API. Streaming and SSE parsing happen in native worker threads, never on the JS loop.
| Export | Behavior |
|---|---|
sofuu.ai.complete(prompt, config?) | Returns a Promise<{ text }> with the full completion. Messages accept an images array (data: URLs) — mapped to OpenAI image_url parts and Anthropic base64 blocks, count/size capped. |
sofuu.ai.stream(prompt, config?) | Returns an AsyncIterator of chunks as they arrive from the model. |
sofuu.ai.embed(text, config?) | Bundled local model by default (SEM2-64, offline, zero-key). Pass an explicit provider/model for network embeddings (BYO key). Array input returns an array; space selects sem2-64 / sem1-64 / hash-768. |
sofuu.ai.embedBatch(texts[], config?) | Local-only vectorized batch — one call, one Float32Array per input. Remote batching lives on ai.embed([...]) with a provider. |
sofuu.ai.embedImage(bytes) | Image bytes (Uint8Array, e.g. from sofuu.fs.readFileBytes) → 64-dim unit vector in img1-64 space, joint with text — so embedLocalSemanticV2 text queries retrieve images. Throws on undecodable input. |
sofuu.ai.transcribe(audio, config?) | Speech-to-text via an OpenAI-compatible audio endpoint (multipart upload, BYO key): Uint8Array or {audio_b64} in, Promise<{ text }> out. No bundled STT — on-device path lives in the SDK wrappers. |
sofuu.ai.speak(text, config?) | Text-to-speech via an OpenAI-compatible endpoint: {voice, format} options (mp3|opus|aac|flac|wav|pcm), resolves {audio: Uint8Array, format}. |
sofuu.ai.embedLocal(text) · embedLocalSemantic(text) · embedLocalSemanticV2(text) | Bundled offline embeddings, sync, no network: 768-dim hash, 64-dim SEM1, 64-dim SEM2. Spaces via embedInfo() / embedInfoV2(). |
sofuu.ai.similarity(a, b) | Cosine similarity of two Float32Arrays on hardware SIMD registers (AVX2 / NEON). |
sofuu.ai.dot(a, b) · sofuu.ai.l2(a, b) | Raw dot product and L2 distance over the same native kernels. |
sofuu.ai.estimate(text) | Fast token-count estimate for budgeting prompts. |
sofuu.ai.model() · sofuu.ai.resolve(name) | Inspect the active model and resolve provider/model names. |
Headless SDK: sofuu_embed_local / batch / image / info and sofuu_voice_transcribe / speak C ABI (libsofuu) with Swift/Kotlin wrappers — same spaces, no JSON. Image recipe: open a store with (path, 64, 'image-projector-v1'),remember image vectors with captions, recall with a text vector.
sofuu.http & fetch
| Export | Behavior |
|---|---|
sofuu.http.createServer(cb) | Binds a non-blocking HTTP server directly over OS sockets. The callback receives req and res. |
server.listen(port, host) | Starts accepting connections. |
res.writeHead(status, hdrs) | Sends the response status and headers. |
res.end(data) | Flushes the body and finishes the response. |
sofuu.http.serve(path) | Mounts a directory of static files onto the server. |
sofuu.fetch(url, opts?) | HTTP(S) client over a statically linked curl — runs on background polling, resolves to a response with .text() / .json(). |
sofuu.fs & processes
| Export | Behavior |
|---|---|
sofuu.fs.readFile(path) | Reads a file from disk. |
sofuu.fs.writeFile(path, data) | Writes (overwrites) a file. |
sofuu.fs.appendFile(path, data) | Appends to a file — handy for local logs. |
sofuu.fs.readdir(path) | Lists a directory as an array. |
sofuu.fs.mkdir(path) · sofuu.fs.rm(path) · sofuu.fs.exists(path) | Directory creation, removal, and existence checks. |
sofuu.spawn() | Launches a child process with an argument list and returns its output — handy for build tools and device automation. |
sofuu.web
Web access for scripts and the agent — the same primitives the assistant uses when you ask it to look something up.
| Export | Behavior |
|---|---|
sofuu.web.search(text) | Searches the web and returns ranked results. |
sofuu.web.open(url) | Fetches a page and returns reader-friendly text. |
sofuu.memory & sofuu.kv
Durable, local persistence backed by the QTSQ encrypted store — the same layer behind chat sessions and long-term memory. Everything stays on your disk, sealed with post-quantum cryptography (ML-KEM-768).
| Export | Behavior |
|---|---|
sofuu.memory.open(path, vecDim) | Opens (or creates) a persistent vector memory at path. |
mem.remember(text) | Stores a memory with an embedding for later recall. |
mem.recall(text) | Returns the most relevant stored memories for a phrase. |
mem.consolidate() | Folds weak and duplicate memories — the decay pass behind the agent's long-term brain. |
sofuu.kv.open(path) | Opens a small encrypted key-value store. |
kv.get · kv.set · kv.keys | Plain persistent key-value access. |
sofuu.tools & sofuu.agent
The agent loop is scriptable: give a model a goal and let it drive the same tool belt the chat uses, with every step emitted as events you can render or log.
| Export | Behavior |
|---|---|
sofuu.tools | The built-in tool belt: read, write, edit, grep, glob, list_dir, bash. |
sofuu.agent.run(task, opts?) | Runs the autonomous agent loop against a task; yields tool calls and results as an event stream. |
Advanced: MCP · RLM · ML
Experimental surfaces — available in the binary but still evolving; expect changes between releases.
| Export | Behavior |
|---|---|
sofuu.mcp.connect(url) | Connects to a Model Context Protocol server and exposes its tools. |
sofuu.rlm | Recursive delegation — splits a hard problem into sandboxed sub-runs. |
sofuu.ml.track(event) | Feeds the on-device advisory models that tune scheduling and compaction (advise-only; never filters your work). |
The five context-economy gates (freshness, compaction, relevance, supervisor, alloc) are tiny offline models — 6–9k parameters each, deterministic, zero API calls. Each is graded on held-out content families against two references: a constant predictor and a logistic regression on the same features. All five beat both.
| Gate | Params | Net F1 | Linear F1 | Constant F1 |
|---|---|---|---|---|
| freshness | 8,721 | 1.000 | 0.571 | 0.000 |
| compaction | 8,201 | 0.850 | 0.553 | 0.545 |
| relevance | 8,617 | 0.989 | 0.484 | 0.000 |
| supervisor | 8,201 | 1.000 | 0.832 | 0.769 |
| alloc | 6,321 | 0.991 | 0.973 | 0.000 |
Compaction is the one irreversible thing the chat does, so it runs under guards that do not trust the model: a per-pass cap (a quarter of turn blocks, two always survive) and a lexical floor that protects any turn where you asked a question, gave an instruction, or recorded a decision. Stated plainly: these gates are trained on synthetic data with mechanical labels, so the table shows they are learnable on the distribution we generate — not that they are correct on your traffic. The honest test is a shadow A/B on real sessions, and we have not built it yet.
Embedding benchmarks
Measured on public datasets, not estimated. Reproduce with scripts/bench/fetch_datasets.sh + ml-train bench. Apple M2 Pro, release build, single thread, no network.
| Space | SciFact nDCG@10 | R@5 | STS ρ | Size / speed |
|---|---|---|---|---|
| BM25 (reference) | 0.663 | 0.750 | — | 0 B |
hash-v1 (default) | 0.410 | 0.493 | 0.613 | 0 params, 0 B, 0.002 ms |
sem1-64 (opt-in) | 0.014 | 0.017 | 0.421 | 13.4k params, 13.7 KB |
sem2-64 (opt-in) | 0.081 | 0.090 | 0.517 | 30k params, 29.9 KB |
The BM25 row is the calibration check: our implementation lands at the ~0.665 published for SciFact, so the harness agrees with the literature. The default embedder reaches 62% of that with zero parameters, zero bytes and no model file — the trade against a hosted embedder is quality for a network round-trip, a key and cents per million documents. The 64-dim learned spaces are in-domain optimizers, not general retrievers; we publish the gap rather than hide it.
Embed the runtime in your app
The CLI above is the free developer surface. To put the runtime inside your own app — a private agent with an encrypted local memory, no cloud account, no API key in your binary, no telemetry — the same engine ships as a library. QuickJS is a pure interpreter, so there is no JIT and the iOS App Store executable-memory policy is satisfied by construction.
| Platform | Add this |
|---|---|
| iOS / macOS | .package(url: "https://github.com/sofuu-runtime/sofuu.git", from: "0.2.0") — or pod 'Sofuu', '~> 0.2' |
| Android | implementation 'com.sofuu:sofuu-android:0.2.0' |
| C / C++ | the libsofuu tarball (sofuu_embed.h + static/shared lib) |
// Package.swift (or: pod 'Sofuu').package(url: "https://github.com/sofuu-runtime/sofuu.git", from: "0.2.0")// App code - a private agent with a memory the user ownsimport Sofuulet sofuu = try Sofuu()let vec = try sofuu.embed("the sky is blue") // 768-dim, offline, no model filelet brain = try Brain.openDefault(sofuu: sofuu) // encrypted, local, yourstry brain.remember("the sky is blue")for hit in try brain.recall("what colour is the sky?", k: 3) {print(hit.text, hit.score)}
Everything above runs on-device. Embeddings ship with no model file at all (0 bytes), the brain is encrypted on disk and exportable, and the full C ABI is behind one header. The full contract, the streaming event schema, and the error model are documented in the repo under docs/EMBEDDING.md.