WUBBERY PUBLIC PROOF NOTES Published: 2026-07-25 WHAT IS PUBLIC - The aggregate benchmark results in results.json and benchmark-results.csv. - Dataset names, metric, run date, seed, baseline scores, module scores and candidate counts for the memory evaluation. - Aggregate request, token, session and cost totals for the private-log cost replay. MEMORY EVALUATION - Harness: BEIR datasets. - Metric: nDCG@10. - Datasets: scifact, nfcorpus and fiqa. - Baseline: dense semantic retrieval with a full scan. - Module result: session-regime candidate selection followed by the same scoring comparison. - Run date: 2026-07-10. - Seed: 7. - Falsification boundary: with the state signal disabled, the module must return to the baseline score. COST REPLAY - Input: the founder's complete private Claude Code usage history used for this reported run. - Aggregate size: 34,734 requests, 9.68 billion tokens and 3,079 sessions. - Comparison: actual logged spend versus a shadow routing decision over the same request history. - Cache pricing was held constant across comparisons. - Raw prompts and session content are not public because they contain private material. IMPORTANT LIMITS - These files make the currently published outputs accessible; they do not make the sealed module implementation public. - The numbers are not a guarantee for a different workload. - A prospective user should request a shadow evaluation using their own traffic and receive their own result before making a commercial decision. - Any claim that cannot be supported by an accessible aggregate output should be treated as provisional, not as a product promise.