Codendum

One machine runs the model. Everyone keeps their code.

Codendum turns a single NVIDIA GB10 into a coding-model service shared by a class or a team. Each person runs OpenCode on their own workstation, and prompts and source code never leave the organization’s network.

Version 0.1.1, released 1 October 2026. Open source under the Apache License 2.0.

How requests reach the modelSeven workstations running OpenCode send requests over HTTPS to an nginx gate, which forwards them to vLLM on a single GB10. Tokens stream back the same way. A host outside the allowed networks is stopped at the gate.OpenCode on each workstationnginx :8443vLLM on one GB10outside the allowed networks

A request makes three stops

Scripts, example configuration and documentation to install the service, secure it, test it, measure it and operate it. Every service runs in Docker, and nothing is installed on the host system.

  1. Workstation

    OpenCode, Git, toolchains, IDEs, builds and tests

    • Each person runs OpenCode on their own machine, with an explicit OpenAI-compatible provider.
    • Code is edited, built and tested here, in any language and for any platform.
    • One command configures OpenCode from the server’s own model name and context limit.
  2. nginx container, port 8443

    The only way in

    • HTTPS from the organization’s LAN or VPN, checked against a network allowlist.
    • One API key per user, with per-user and global limits.
    • Only /v1/chat/completions and /v1/models are forwarded. Everything else stays on the host.
  3. vLLM container on the GB10

    Qwen3-Coder-30B-A3B-Instruct-FP8, served as “coder”

    • Listens on 127.0.0.1 only, with tool calling enabled.
    • FP8 KV cache, chunked prefill and prefix caching.
    • The GB10 only serves inference: CPU and GPU share the same memory, so builds run elsewhere.

Measured on a GB10, not promised

40 simulated students ran OpenCode-like agent sessions on a Java project against one Lenovo ThinkStation PGX, through a WireGuard VPN with about 100 ms of round trip. Each made four requests with 20–90 s pauses.

2K and V×48layers×4KV heads×128head size×1 byteFP8=48 KiBper token

After 31.2 GB of FP8 weights, vLLM reported 62.4 GiB of KV cache: 1,362,640 tokens, or 20.8 full 64K contexts at the same time.

Classroom simulation results for two serving profiles
40 simulated studentsclassroom-64k16 running…-high-concurrency24 running
Server output, steady state~110 tok/s~140–150 tok/s
Time to first token, P50 / P9517.5 / 60 s10.6 / 34 s
Output speed per user, P507.2 tok/s6.1 tok/s
Time to complete a coding request, P50 / P9511 / 22 min8.3 / 18 min
Coding requests completed4564
Peak KV cache usage8.5 %11.5 %
Preemptions / server errors0 / 00 / 0
Prefix-cache hit rate98.2 %98.2 %

The limit is generation, not memory

Agent contexts stayed between 7K and 19K tokens, so the KV cache never went above 12 %. Admitting more requests at once raised total output.

Prefix caching does the heavy lifting

98 % of the 7 million prompt tokens were served from the cache.

Plan in coding requests per hour

About 2,700 output tokens over 15 model calls each: roughly 150–200 requests per hour for the whole class.

Full method and scripts in the benchmark documentation. Run bench.sh and bench-classroom.sh on your own host.

From a fresh GB10 to a working class

On the host, with Docker and the NVIDIA runtime that DGX OS provides. The installation guide and the proxy guide explain every step.

  1. Check the host

    Read-only checks: ARM64, Docker, NVIDIA runtime, memory, disk, port exposure, image pinning.

    git clone https://github.com/stefanoferi/codendum.git
    cd codendum
    scripts/preflight.sh
  2. Start the model

    vLLM binds to loopback only and refuses unpinned images.

    cp .env.example .env
    scripts/start-vllm.sh --wait
  3. Open the gate

    Set the host name and allowed networks, install a TLS certificate and generate per-user keys, then start the proxy.

    scripts/start-proxy.sh --init
    scripts/gen-api-keys.sh --users-file users.txt
    scripts/start-proxy.sh
  4. Test it

    Health, model list, chat, streaming, tool calls and a tool-result round trip.

    scripts/smoke-test.sh
  5. Connect each workstation

    Writes the OpenCode provider from the server’s limits, runs a test completion and backs up the previous file.

    scripts/configure-opencode.sh --base-url https://llm.lab.example:8443

What Codendum does not do

The limits are written down, so you can decide whether they matter for your class or your team before you install anything.

One host, no high availability
When the GB10 or vLLM is down, every user is affected.
No per-user token quotas
nginx caps requests, not tokens, and vLLM serves requests first come, first served.
The prefix cache is shared
Timing could in principle reveal that someone recently sent the same prefix. You can turn prefix caching off, at a throughput cost.
Agents run on the workstations
OpenCode executes tools with the user’s permissions. Sandboxing belongs on the workstation side.
Policies are yours
Prompts contain users’ code and the logs contain user ids. Data protection and log retention are the operator’s responsibility.
Numbers are for this setup
One GB10, one model, one client version. Measure again with the benchmark scripts after every upgrade.

Read the full threat model and report vulnerabilities privately as described in SECURITY.md.

Next: governing what goes through the model

Today Codendum controls who can use the model. A possible next step is governing what goes through it: personal-data redaction, a prompt-injection firewall, agent loop breaking and a tamper-evident audit log.

That could come from Admina, an open-source framework for governed AI, placed as a gateway between the proxy and vLLM. It is a working hypothesis: there is no code and no timeline yet. See Governance.