Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
-
Updated
Jul 22, 2026 - Python
Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
A negative result on joint-embedding predictive architectures for prediction markets. The martingale-collapse diagnostic is sound and predicts nothing, more training makes the representation worse, and the metric everything was ranked by was mostly measuring input reconstruction. Pre-registered gates, full findings record.
JEPA agent playing Minecraft from pixels: latent world model + MPC planning, 664K params on one 8GB GPU, trained on raw gameplay with no labels. A complete lab notebook - including a 20-attempt research dead end, documented with its root cause.
[zenodo.20574533] CGAA: Concept Guided Adversarial Attacks
Distribution-shift teardown of openpilot v0.9.7 supercombo: does a production L2 self-driving model know when it's blind? (it doesn't, and silently)
"Publication bias and the canonization of false facts" published in eLife (2016)
Reproducible evaluation harness for hidden coordination variables in multi-agent LLM systems.
A $0 falsification lab across two markets — crypto (~111 hypotheses) and Polymarket prediction markets (172,830 resolved markets, 1.36M trades). 184+ techniques through one committed anti-overfitting gauntlet. 0 survive to a deployable edge; the reusable validation harness is the asset. Agent-ready. MIT.
从20个项目中系统提取可复用方法论模式的实验记录。包含10轮正式审查实证(4后端)、58项发现、G5可追溯审计。不成熟框架,诚实的实验记录。
Measurement-first research on local Mixture-of-Experts inference under a hardware contract. 3 measured laws, 4 falsified ideas, and paper site.
Research platform for discovering and rigorously falsifying crypto trading strategies: event-sourced paper trading, implementation-parity verification, and a documented negative-results record.
Signed attention for transformers — sinh/cosh attention giving weights in [-1,1] instead of softmax's positive-only, made FlashAttention/SDPA-compatible by channel doubling. Includes a from-scratch softmax-vs-SBA comparison and a documented negative result.
A rigorous negative result: no pre-publish feature predicts YouTube Shorts engagement above chance, and the one 95% model is a leakage trap.
Synchronization Resistance: a pre-registered study measuring multi-LLM agreement difficulty. 3 pass / 1 inconclusive / 1 fail — critique wanted.
A global, open-source registry and standardized schema for null findings, non-significant outcomes, and failed trials in medical research. Built to eliminate publication bias and accelerate biomedical discovery.
A negative-result study and a falsification protocol for LLM memory systems.
Component-level benchmark for catastrophic forgetting in world models. Two negative results: forgetting does not follow the labelled task-distance axis, and it happens in the encoder, where the usual metrics cannot see it. 375 runs, with code, data and paper.
Reproducible Apple Silicon benchmark: adaptive 2/4-token prompt lookup did not beat fixed-2 on Qwen3-0.6B.
Open-source market intelligence platform with a self-auditing research pipeline: pre-registered trials, placebo gates, published negative results, live paper track record vs SPY. FastAPI + Next.js + LightGBM.
Diagnostic study of adaptive gating failure in vision-language prompt learning, focusing on gradient imbalance and gate collapse under frozen CLIP backbones.
Add a description, image, and links to the negative-results topic page so that developers can more easily learn about it.
To associate your repository with the negative-results topic, visit your repo's landing page and select "manage topics."