Open research toward agents that run unattended and models that are cheap to run and easy to inspect.
TBC Research builds open-source tooling for long-running Claude agents and publishes reproducible research on efficient inference and model diagnostics, including the negative results.
Three research areas
Frontier models can now work for hours, but most tooling still assumes a person watching the terminal. We work on the pieces that close that gap, and on the systems research underneath them.
Agent infrastructure
Bounded, verifiable, long-running Claude Code sessions, and procedural memory that persists across them.
Efficient inference
From-scratch inference engines and kernels for Apple Silicon and commodity GPUs, measured against llama.cpp and MLX.
Model diagnostics & representation
Checkpoint lifecycle diagnostics, auditable compiled decoders, and compositional representations.
All public repositories
Every project below is public on GitHub. Descriptions are summarised from each repository's README.
An autonomous overnight loop for Claude Code
Splits a mission into fresh-context Claude Code episodes on Opus, with structured handoffs, per-episode budget caps, git-diff checks of claimed progress, and eight automatic stop conditions. A research toolkit monitors arXiv and Semantic Scholar and synthesises findings.
Hebbian procedural memory for coding agents
Learns which files are accessed together, which errors lead to which fixes, and which tool chains recur, then recalls them before the agent searches. Installs into Claude Code hooks and MCP with npm install brainbox-hebbian. No vector database.
Hybrid Neural Engine, GPU, and CPU LLM inference on Apple Silicon
A from-scratch engine that runs Qwen3.5-2B across the Apple Neural Engine (prefill via private APIs), Metal GPU (decode with custom shaders), and CPU. Metal decode matches llama.cpp at about 32 tok/s. No CoreML, Python, or MLX.
A batch-1 INT4 decode megakernel on a 2018 Turing GPU
A GPTQ INT4 decode megakernel on a Quadro RTX 4000 that runs 1.09–1.64× faster than llama.cpp Q4_0 on Qwen3-0.6B and 1.7B. The abandoned lattice-coding approach and every other dead end are documented.
Speculative speculative decoding on an M4: a documented ceiling
Tests whether SSD-style parallel speculation can beat plain MLX speculative decoding on Apple Silicon. Answer: no. 66.9 tok/s is the practical ceiling, set by memory bandwidth. Along the way it measures real ANE and GPU parallelism and a fast CoreML-to-MLX KV-cache handoff.
A real-time neural synthesizer on the Apple Neural Engine
Runs neural audio inference directly on the Neural Engine, bypassing CoreML: about 157 µs per 8-voice buffer, roughly 79× real-time headroom, with no CPU cores spent on inference.
Measuring and repairing a checkpoint's “surgery reserve”
Shows that two transformers with identical outputs can respond very differently to quantization, LoRA, or merging, demonstrates the effect on real Qwen checkpoints, and provides the tooling to detect and repair it, backed by an append-only evidence ledger.
Self-stabilising, auditable brain-computer-interface decoders
Treats decoders as programs: diagnoses where they drift, repairs them from about 30 s of data with label-free correction on human intracortical data (FALCON H1), compiles them to dependency-free streaming Rust, and certifies the algorithm small RNNs actually run. Preprint: DOI 10.5281/zenodo.19339860.
A compositional morpheme tokenizer built on holographic reduced representations
Turns each word into one fixed-width vector by binding prefix, root, and suffix with circular convolution and unitary roles, so morphemes can be algebraically recovered and related words land near each other. It is an embedding layer for experimental models, not a drop-in replacement tokenizer.
How we work
- Open by default. Code, benchmarks, and reports are public under MIT.
- Negative results count. Ceilings and dead ends are written up next to the wins.
- Verify, don't trust. Agent progress is checked against git; claims are tied to evidence.
- Measure on real hardware. Every number is taken from a named device and a named baseline.
Get in touch
Research collaborations, feedback on the tools, or notes on agent reliability and efficient inference are all welcome.
bb@tbcresearch.org