WUBBERY SUBSTRATE ENGINE & COMPANY / SUB-2.35μs ENTERPRISE ROUTER

Intelligence in motion.
WUBBERY is the world's first sub-2.35μs AI Substrate Engine. Powered by the Harkin Theorem (8 patent families), WUBBERY sits between applications and AI models to route 425,000 decisions/sec, slash data centre megawatt power draw by 88%, and eliminate token cost waste on CPUs and GPUs.
Let's talk about our future
Any data centre not routing like this — do they care about the Earth?
The Data Centre Grid Crisis vs WUBBERY Substrate Salvation
A two-act comparison of how unrouted cloud AI scorches electrical grids vs how WUBBERY on-device WASM preserves Earth.

Data Centre Grid Scarcity & Latency Lag
Unrouted AI models send simple queries to distant GPU server farms. Every request consumes megawatts of cooling electricity, adds 2,400ms of lag, and inflates cloud bills.

On-Device Zero-Marginal-Cost Execution
WUBBERY runs WebAssembly bytecode natively on the user's browser or edge CPU. Simple intent queries resolve in 2.35μs with 0% cloud server overhead, keeping Earth green.
Initial module set
Eight modules first.
Fifteen layers all up.
per routing decision
425,000 decisions/sec on one CPU thread. Other routers spend an AI call to route AI calls. This is infrastructure.
on 34,734 real requests
9.68B tokens of the founder’s own usage, shadow-replayed and priced three ways. Nothing simulated.
memory lift · ambiguous regime
BEIR fiqa, the field’s own metric — while scanning 19.7× fewer candidates.
within-one-tier accuracy
Held-out routing suite. Misses printed, not hidden. Re-runnable in front of you.
predictive prefetch hit-rate
On structured usage — and ~3% on random. We publish where it does NOT help.
inference energy
Versus all-premium baselines. The routing dividend, measured in watts as well as dollars.
modular layers · released in stages
WUBBERY is a 15-layer modular platform. The initial release brings the six modules shown here, plus Governance and Anomaly Detection. The remaining seven layers will be introduced in stages.
Against the field
The industry's best, at their best. Then ours.
Test V/01 · published vs live-measured
Decision latency per request · lower is better
EV · Embedding semantic router
5msthe field’s "fast" tier — impressive — five thousandths of a second
RL · RouteLLM classifier
84msthe published state of the art — a model routing models — research-grade
ND · Not Diamond
100msthe venture-backed leader — "worthwhile tradeoff for the savings"
W · ROUTING MODULE — measured live, one CPU thread
0.00235ms2.35 microseconds. The bar is three pixels wide because it has to be visible at all.
Bars drawn on a log scale, or every competitor would be off the page.
Now ours
faster than the published state of the art
42,553× faster than the leader's best case · 2,127× faster than the field's "fast" tier.
Test V/02 · same numbers, true scale
At true scale our bar is thinner than anything this screen can draw. That's what infrastructure looks like next to an AI call. Theirs need a GPU-backed service; ours runs air-gapped on anything.
Fair fight: their figures are their own published/reported decision overheads; ours is measured live on a laptop CPU — run GET /v1/bench/overhead and watch your own number.
Memory · public data
Three benchmarks. Three domains. The lift scales with ambiguity.
SCIFACT
science · unambiguous
13.6× less workNFCORPUS
medicine · hard
12.4× less workFIQA
finance · ambiguous
19.7× less workFalsify this
The boundary, published on purpose: with no state signal the routing module gates to zero influence — identical to baseline to four decimal places, on all three datasets. Set β=0 and the lift must vanish. Set noise=1.0 and the routing module must equal baseline. It does.
Test B/01–03 · third-party harness · run 2026-07-10 · seed 7 · deterministic
Cost · real traffic
One machine. Every request. Three worlds, priced.
Test S/01 · live-measured logs
34,734 requests · 9.68B tokens · 3,079 sessions
All-premium · naive ceiling
$14,792Actual spend · as logged
$13,017Module-routed · shadow
$8,734module-routed vs actual spend, on the same 34,734 requests
Test S/02 · pick what you run today
Premium
Fable-5 / GPT-5-class
Balanced
Sonnet-5 / GPT-4o-class
Budget
Haiku / Llama-class
−42.4%
projected cost
89.8%
of requests
94.4%
tier-hold accuracy
Hard requests keep the best model — the saving never comes out of quality.
Coding agents · Fable 5 baseline, priced both ways
Claude Code on Fable 5. Same session, 80.7% cheaper.
Without routing · every turn on Claude Fable 5
$6.41With routing · right-sized model per turn
$1.24The top-tier baseline, not a strawman — hard turns stay capable, no quality cliff.
One session · every turn, routed
An Anthropic request never routes off-family. Per-model prices are representative published tier rates — the dollars are illustrative; the tier mix and the rate are measured. Prove it on your own prompts at zero spend with X-LEWM-Dry-Run: 1.
What this means for you
Six dials, translated into your world.
2.35µs
Your users never feel the router.
Every other approach adds 5–200ms to every single request. Ours adds nothing a human — or an SLA — can detect.
−32.9%
A third of your AI bill, back.
Measured on real mixed traffic, not a marketing mix. Hard requests still get the best model.
+61.2%
Better answers where it’s hardest.
Ambiguous questions — the ones your support and search actually struggle with — is exactly where the memory lift is largest.
94.4%
Savings without the quality tax.
Routing accuracy means the cheap model only gets what the cheap model can do. That is the whole trick, held to a published number.
94%
Feels faster than it has any right to.
Prefetch means the answer is often warming up before the request lands — 2.7× faster with use, on structured workloads.
−71%
Your GPUs go 3.4× further.
If you run in-house models, energy IS capacity: the same rack serves multiples of the traffic — or the same traffic on a third of the power.
and that's only what we show. Eight patent-pending families, each holding more than one claim. The rest is ledgered too — it just needs an NDA to see.
Modular by design
Choose the layers your business needs.
ROUTE
2.35µs
RECALL
+61.2%
PREFETCH
94% warm
TIER HELD
94.4%
LEDGER
−32.9%
ENERGY
−71%
GOVERNANCE
INITIAL
ANOMALY
INITIAL
Because the platform is modular, a business can start with the layers it needs and add more over time. When selected modules run together, the right answer can be ready before the request arrives — 100% of the time in our ledgered test, versus 0% when the same capabilities run as disconnected point-tools. Customers get a staged path into the platform without giving up the compounding benefit.
How it actually plugs in
No rewrite. No SDK. One address change.
Watch first — shadow mode.
Zero risk, zero code. We replay your existing logs and show the ledger of what you would have saved. This is the free tier, and it’s how every engagement starts.
$ replay --dir your-logs/— your number, zero spend
Route live — change one address.
Every app that calls AI has a setting for where to send its requests. Today it points at your provider. Change that one setting to the routing module’s local address — your keys, your providers, your data path. One settings line, no code.
$ OPENAI_BASE_URL=http://localhost:8787/v1— that’s the whole integration
Ask before you spend — the decision API.
Keep full control: your code asks "which model should this request use?" and gets an answer in microseconds, then makes its own call. For in-house fleets, this turns −71% energy into 3.4× capacity.
$ POST /v1/route {"text": "..."}— sealed decision, nothing leaves your machine
What a company rollout actually looks like · one engineer · hours, not months
Shadow your history.
Your platform team exports existing AI logs. We replay and hand back a ledger: what you spent vs what you would have. Nothing installed, nothing leaves your side.
One workload goes live.
Pick something low-stakes — an internal tool, a summarisation job. One engineer changes its provider address. The ledger starts measuring for real. Rollback is changing the address back.
Expand team by team.
Each workload is its own one-line switch, each shows its own ledger. Security reads the audit log; finance reads the savings. Stop any workload any time.
Inside your walls.
For regulated environments selected WUBBERY modules install on your own servers — your IT, your network rules, air-gapped if needed. Per-workload calibration and sealed capabilities come with the NDA engagement.
GET YOUR NUMBER
Shadow access is free and private — tell us roughly what your traffic looks like and we'll set you up. No card, no commitment, no data leaves your side.
Request module access →Security & privacy
Built like it's guarding something. Because it is.
Your prompts stay yours
Decisions happen on your machine in microseconds. In decision mode nothing leaves your network. In gateway mode requests go to your providers with your keys — the platform stores none of it.
No phone home
The service binds to localhost and refuses the network by default. Exposing it is an explicit operator act that hard-requires its own token — and TLS in front.
Keys pass through, never stored
Provider credentials ride the request and are forwarded only to the fixed upstream you configured. No key ever touches disk.
Every decision auditable
An append-only audit log records every route, gateway call, and failed auth — outcomes only, timestamped, safe to ship to your SIEM.
Hardened by default
Timing-safe token checks, brute-force lockout (10 failures → 5-minute ban), per-tier rate limits, request-size caps. On by default, not opt-in.
Nothing to supply-chain
The sealed service has zero third-party dependencies — no packages to poison, no CVE surface to babysit. And it runs fully air-gapped: banks, health, and government keep everything inside the perimeter.
The bank case — your machines, improved
Regulated industries can't ship prompts to a routing API — so no other router can even bid. WUBBERY installs on your own hardware, inside your own perimeter, and makes the machines you already own better: 3.4× effective capacity from the energy dial, answers 2.7× faster with use, a third off any external spend — with the audit log your compliance team already wants. Nothing new to trust. Your metal, upgraded.
Badges marked pathway are in progress, not held — certification lands with the first enterprise pilots. The architecture was built for their audits from day one.
The seal
The seal works both ways: responses carry the model, tier, and outcome scores — never the internal state that produced them. Reverse-engineering the service tells you what it decided, not how. Your data can't leak what we never transmit.
At scale
For data centres and big business, this is an energy story.
For data centres
Your bottleneck is power, not silicon. We hand you power back.
- 3.4× the inference from the same grid connection. −71% inference energy isn't a saving line — it's headroom. The megawatts you already have suddenly serve more than three times the billable load.
- Cooling follows compute. Less wasted compute means less heat: PUE improves without touching the plant room.
- Deferred capex. The next hall, the next substation, the next grid negotiation — pushed out years, because the building you have does more.
- Grid-constrained sites win. Where you can't get more power, efficiency is the only growth path. This is that path.
For big business
A third of the AI bill, back on the P&L — not on a roadmap.
- Measured, not promised. The same ledger you watched in shadow becomes the invoice basis — savings are counted from your own traffic before anyone is paid.
- Procurement leverage. When any workload can take any model, no single provider owns you. Negotiate from the ledger.
- Sovereignty kept. Runs inside your perimeter, air-gapped if required — regulators, auditors, and security teams all read the same audit log.
- Capacity planning flips. Buy hardware for today's real load, not 3.4× of it. Growth rides the efficiency curve instead of the purchase order.
−71%
inference energy — off the meter, at fleet scale
×3.4
effective capacity from the hardware you already own
0
new racks, new grid connections, or new code required
100%
auditable — every routed decision, timestamped
Your number
What would it mean for you? Move the dials.
GPUs serving inference
1 GPU → serves like 3
What that means for a single GPU
One card suddenly serves the traffic of three and a half — the energy dial turns −71% inference energy straight into headroom. No new hardware, no bigger power bill: the same GPU you already own takes 3.4× the users, or the same users with the fans barely spinning. And routing costs it nothing — decisions run on the CPU in microseconds, so the GPU only ever does model work.
×3.4
effective capacity — −71% inference energy: same racks, same power budget, or the same load on a third of the power
2.7×
faster with use — measurably self-improving on structured workloads: the longer it runs, the quicker it gets
0 µs
GPU spent on routing — decisions run on CPU in microseconds: the router never queues behind your models
Projected from measured rates on our ledgered runs. Your mix will differ — shadow mode replays your own logs and shows your true number before you spend a cent.
Module Access
Request the layers your business needs first.
Request
pricing confirmed with scope
Modules and release group
Routing · Memory · Prefetch · Tiering · Ledger · Energy
- ✓Choose individual modules
- ✓Discuss your workload and environment
- ✓Confirm availability before onboarding
- ✓No combined-OS requirement
Request
pricing confirmed with scope
Modules and release group
Governance · Anomaly Detection
- ✓Scoped permissions and policy controls
- ✓Approval and audit requirements
- ✓Operational anomaly visibility
- ✓Security requirements assessed first
Register
interest in the next releases
Modules and release group
Seven further layers · Fifteen total
- ✓Tell us what capability is missing
- ✓Join the relevant release list
- ✓Request more than one module
- ✓No purchase or account commitment
Accessible proof
Don't take our word for it. Open the outputs.
— structured aggregate outputs
— every published benchmark row
— scope, method, privacy limits and caveats
Ready when you are
Start with shadow — free, private, your own logs, your true number. Or come straight to an NDA briefing for the sealed dials.
The rest of WUBBERY
