One machine runs the model. Everyone keeps their code.
Codendum turns a single NVIDIA GB10 into a coding-model service shared by a class or a team. Each person runs OpenCode on their own workstation, and prompts and source code never leave the organization’s network.
A request makes three stops
Scripts, example configuration and documentation to install the service, secure it, test it, measure it and operate it. Every service runs in Docker, and nothing is installed on the host system.
Workstation
OpenCode, Git, toolchains, IDEs, builds and tests
- Each person runs OpenCode on their own machine, with an explicit OpenAI-compatible provider.
- Code is edited, built and tested here, in any language and for any platform.
- One command configures OpenCode from the server’s own model name and context limit.
nginx container, port 8443
The only way in
- HTTPS from the organization’s LAN or VPN, checked against a network allowlist.
- One API key per user, with per-user and global limits.
- Only /v1/chat/completions and /v1/models are forwarded. Everything else stays on the host.
vLLM container on the GB10
Qwen3-Coder-30B-A3B-Instruct-FP8, served as “coder”
- Listens on 127.0.0.1 only, with tool calling enabled.
- FP8 KV cache, chunked prefill and prefix caching.
- The GB10 only serves inference: CPU and GPU share the same memory, so builds run elsewhere.
Measured on a GB10, not promised
40 simulated students ran OpenCode-like agent sessions on a Java project against one Lenovo ThinkStation PGX, through a WireGuard VPN with about 100 ms of round trip. Each made four requests with 20–90 s pauses.
2K and V×48layers×4KV heads×128head size×1 byteFP8=48 KiBper token
After 31.2 GB of FP8 weights, vLLM reported 62.4 GiB of KV cache: 1,362,640 tokens, or 20.8 full 64K contexts at the same time.
| 40 simulated students | classroom-64k16 running | …-high-concurrency24 running |
|---|---|---|
| Server output, steady state | ~110 tok/s | ~140–150 tok/s |
| Time to first token, P50 / P95 | 17.5 / 60 s | 10.6 / 34 s |
| Output speed per user, P50 | 7.2 tok/s | 6.1 tok/s |
| Time to complete a coding request, P50 / P95 | 11 / 22 min | 8.3 / 18 min |
| Coding requests completed | 45 | 64 |
| Peak KV cache usage | 8.5 % | 11.5 % |
| Preemptions / server errors | 0 / 0 | 0 / 0 |
| Prefix-cache hit rate | 98.2 % | 98.2 % |
The limit is generation, not memory
Agent contexts stayed between 7K and 19K tokens, so the KV cache never went above 12 %. Admitting more requests at once raised total output.
Prefix caching does the heavy lifting
98 % of the 7 million prompt tokens were served from the cache.
Plan in coding requests per hour
About 2,700 output tokens over 15 model calls each: roughly 150–200 requests per hour for the whole class.
Full method and scripts in the benchmark documentation. Run bench.sh and bench-classroom.sh on your own host.
From a fresh GB10 to a working class
On the host, with Docker and the NVIDIA runtime that DGX OS provides. The installation guide and the proxy guide explain every step.
Check the host
Read-only checks: ARM64, Docker, NVIDIA runtime, memory, disk, port exposure, image pinning.
git clone https://github.com/stefanoferi/codendum.git cd codendum scripts/preflight.shStart the model
vLLM binds to loopback only and refuses unpinned images.
cp .env.example .env scripts/start-vllm.sh --waitOpen the gate
Set the host name and allowed networks, install a TLS certificate and generate per-user keys, then start the proxy.
scripts/start-proxy.sh --init scripts/gen-api-keys.sh --users-file users.txt scripts/start-proxy.shTest it
Health, model list, chat, streaming, tool calls and a tool-result round trip.
scripts/smoke-test.shConnect each workstation
Writes the OpenCode provider from the server’s limits, runs a test completion and backs up the previous file.
scripts/configure-opencode.sh --base-url https://llm.lab.example:8443
What Codendum does not do
The limits are written down, so you can decide whether they matter for your class or your team before you install anything.
- One host, no high availability
- When the GB10 or vLLM is down, every user is affected.
- No per-user token quotas
- nginx caps requests, not tokens, and vLLM serves requests first come, first served.
- The prefix cache is shared
- Timing could in principle reveal that someone recently sent the same prefix. You can turn prefix caching off, at a throughput cost.
- Agents run on the workstations
- OpenCode executes tools with the user’s permissions. Sandboxing belongs on the workstation side.
- Policies are yours
- Prompts contain users’ code and the logs contain user ids. Data protection and log retention are the operator’s responsibility.
- Numbers are for this setup
- One GB10, one model, one client version. Measure again with the benchmark scripts after every upgrade.
Read the full threat model and report vulnerabilities privately as described in SECURITY.md.
Next: governing what goes through the model
Today Codendum controls who can use the model. A possible next step is governing what goes through it: personal-data redaction, a prompt-injection firewall, agent loop breaking and a tamper-evident audit log.
That could come from Admina, an open-source framework for governed AI, placed as a gateway between the proxy and vLLM. It is a working hypothesis: there is no code and no timeline yet. See Governance.
