We build the evidence layer for AI-written code.
Coding agents know your code. They have never understood how it behaves when it runs. Vinv records the run, joins every call back to the exact line that served it, and hands that graph to your agent — then judges every fix on what the code actually did.
Free · open source · Apache 2.0 · 100% local, no telemetry, no accounts, no API keys.
vinv.ai · Install · Open VSX · LinkedIn · support@vinv.ai
Software that an AI writes should be judged by what it does, not by what it claims.
84% of developers now use or plan to use AI coding tools — and more of them actively distrust the output (46%) than trust it (33%), with distrust nearly doubling in a year (Stack Overflow 2025, 49k developers). Everyone knows the failure shapes: the agent edits the wrong handler, invents a return type, then grades its own homework while the server won't even start. Or it enters the doom loop — same failed edit, same error, over and over, burning the context window on "let me verify."
Both have one root cause: the agent has never watched your code run. It argues from static text.
So we're building the missing half of the loop. Not a better model, not another prompt: a runtime evidence layer that watches real executions, ties every span to the symbol that produced it, and stands between "the agent says it's fixed" and "it's merged." Our bet, and the thing every release has to keep proving: context beats model size.
The end state we're working toward — a developer hands work to an agent and gets back not a diff and a promise, but a diff, the run that exercised it, the tests it never saw, and a verdict that holds up when someone checks.
Vinv — one product, nine engines, installed as an editor extension and driven by the coding agent you already pay for. It runs your service, exercises it, finds what's broken or slow, dispatches the fix to your agent, and verifies the result independently.
| Runtime tracing | Zero-edit tracer — no SDK, no decorators, no code changes. Timing, memory, arguments, returns and errors, per call. |
| Code Graph + semantic search | A Rust semantic index and call graph over every symbol, embedded by a model that runs on your machine. Ask by meaning, get ranked symbols with line numbers. |
| The context graph | Code, traces and the metrics derived from them — joined on the exact function that handled each request. The artefacts are commodities; the join is not. |
| Behavior exerciser | Doesn't wait for traffic. Drives every endpoint itself with schema, boundary, negative and multi-step auth inputs, and banks every response as a permanent regression case. |
| Verified fixes | Replayed start, live port, and acceptance tests written before the fix that the agent never sees. One click reverts everything an episode touched. |
| Optimization with statistics | A speedup lands only if a paired-bootstrap 95% CI excludes zero and behavior replays byte-identical. Faster-but-wrong is auto-reverted. |
| Agent babysitting | A doom-loop guard catches a repeating agent, an adaptive watchdog catches a hung one, and "Dispute a Verified Fix" keeps the verifier accountable too. |
| MCP servers | vinv-index and vinv-runtime serve the evidence to Claude Code, Cursor, Codex CLI, Copilot Chat, Windsurf and Gemini CLI. |
Editors: VS Code · Cursor · Windsurf · VSCodium · Trae · VS Code Insiders
Stacks: Python backends get the full loop today. Other languages get the index, graph and grounded Q&A — runtime evidence for TypeScript and Go is next. We say this out loud rather than burying it.
Five real issues in fastapi/full-stack-fastapi-template (44k★) — four bugs and one performance problem, all found by Vinv itself. Same issues, same prompts, Vinv grading every run:
| Setup | Fixed |
|---|---|
| Cheap commodity model + Vinv context | 4 bugs + 1 optimization |
| Frontier model, working blind | 1 bug |
| Cheap commodity model, working blind | nothing |
A demonstration, not a benchmark — five issues, one repo, one trial per condition. Published because it's checkable, not because n=5 settles anything.
The claim isn't a model ranking. Blind, the commodity model scored zero. A model holding the failing frame, the caller chain and the real argument values beats a stronger model guessing from static code. The evidence is what moved, not the weights.
On the same pristine template the loop later found something nobody planted: the default database pool makes requests queue for connection checkouts under load. Fix dispatched, then proved — sustained-load median 75.6ms → 41.2ms, 45.4% faster, 95% CI [36.3%, 45.8%], responses byte-identical. Two earlier attempts whose measurement windows couldn't certify the win were auto-reverted. The accept landed only when the evidence did.
Evidence over assertion. Every claim we ship traces to something observable — a span, a line, a confidence interval. "Should be faster now" is not a result.
Nobody grades their own homework — including us. Acceptance tests are authored before the fix and hidden from the agent. Retrieval changes are promoted only through off-policy evaluation gates; to date those gates have declined every candidate we proposed. Vinv's own release gate is Vinv.
Honest scope beats a bigger claim. Python first, other stacks partial, and we write that in the README instead of the footnotes. If a number came from one run, we say so.
Your machine, your code. Everything runs locally — per-repo state in .vinv/, per-machine in
~/.vinv/. No account, no provider keys, no telemetry, none. Traces store bounded summaries,
and sensitive parameter names are redacted rather than captured. The only LLM Vinv talks to is the
agent CLI you configured, through your own auth.
Open by default. Apache 2.0, every engine building from source, in one public repository. The algorithms behind each decision are named — Thompson sampling, paired bootstrap, Theil–Sen, Daikon-style invariants, Nash bargaining, spectrum-based fault localization — so a skeptic can check the method, not just the marketing.
Respect the developer's budget. Tokens, attention and trust are all finite. A bounded run that ends in a verdict beats an unbounded one that ends in a bonfire.
| VinvAI/VinvAI | The whole thing — Python engines, Rust index, editor extension, docs. Apache 2.0. |
| Install | One click, pick your editor — or code --install-extension VinvAI.VinvAI |
| Open VSX listing | Marketplace page and reviews |
| vinv.ai | The short version, with the clips |
git clone https://github.com/VinvAI/VinvAI ~/.vinv/engines && cd ~/.vinv/engines && ./install.shFirst run builds the engines (~4 min: compiles the Rust index, fetches a one-time ~500 MB local embedding model). Needs uv and Rust.
Issues, discussions and PRs are open — good first issues are labeled. Start with CONTRIBUTING.md; agents contributing on your behalf should read AGENTS.md. By taking part you agree to our Code of Conduct; to report a vulnerability, see SECURITY.md.
If Vinv caught something your agent missed, tell us — a review on Open VSX or a ⭐ on the repo is how this reaches the next developer stuck in a doom loop.
vinv.ai · GitHub · Open VSX · LinkedIn · support@vinv.ai
© 2026 VinvAI · Apache License 2.0 · Context beats model size.
