The problem
An agent working on a real project does not see the whole project. It sees what it searches for, and it searches for what occurs to it. When the helper it needs lives in another module under another name, it writes it again. When a task allows a short solution and a long one, nothing pushes it toward the short one.
The usual answer is written rules: a CLAUDE.md, a skill, the system prompt. They help, but they are advice. The model may not load them, may forget them halfway through a turn, or may decide they do not apply. We cannot touch the model's weights, but we can control three things: what the model sees, what it is allowed to do, and when it is allowed to finish. Sens uses all three at once.
The circuit
Claude Code accepts hooks as callbacks on its own stream protocol. Sens answers five of them inside its process, with one invariant: a change is born unapproved. Only an audit that passes, or a person, approves it.
- 01
You send a message
UserPromptSubmit
Sens takes a checkpoint of the project and hands Claude up to eight things that already exist and relate to the request, each with its signature, file:line and number of uses.
- 02
Before every write
PreToolUse · Write, Edit
Sens rebuilds the resulting file in memory and applies the change rules. It can deny the write before it happens, with the reason and the code to reuse.
- 03
Before every command
PreToolUse · Bash, PowerShell
Sens protects its own paths and the configuration, and treats git commit and git push as the end of a turn: nothing unapproved gets committed.
- 04
After every tool
PostToolUse
Whatever the tool, Sens looks for the files that changed on disk and applies the same rules to them.
- 05
When the turn ends
Stop · SubagentStop
Sens audits the whole diff since the last approved point: change rules, orphaned code and, if code changed, the reviewer. With blocking findings Claude keeps working; after three rounds the turn is held for you.
Nothing to go around
Before building anything, we checked live that the mechanism holds against a model trying to slip past it: a disableAllHooks written mid-session does not stop the callbacks, subagents go through them, and a denied git commit leaves the repository as it was. Every shortcut has something that closes it:
| Shortcut | What closes it |
|---|---|
| Writing through the shell, Python or another MCP server | The check of the disk after every tool, and the audit when the turn ends |
| Subagents and background tasks | They go through the same hooks; the turn is not approved while they run |
| Turning the hooks off or editing the configuration | The hooks live in Sens's process; rule R7 blocks and restores those files |
| Opening a worktree of its own | Worktree tools are denied, and R7 covers git worktree |
| Declaring the work finished | The audit decides when a turn ends, not the model |
| Starting a new turn to escape | The audit runs from the last approved point, so whatever is pending carries over |
| Committing unapproved work | A commit is treated as the end of a turn |
| Looping forever | Three rounds, then the turn is held |
The limit is explicit: the circuit closes the mistakes and shortcuts of a model, not those of a hostile program running on the same machine.
Results
Single tasks start from a clean project, and the harm of an agent that does not reuse is not in one task but in the sum. Horizonte measures the sum: a TypeScript command line for household expenses that starts at 88 lines, and 30 product requests in a fixed order, each building on the last. Eight concepts are planted that several tasks need without saying so: dates, months and weeks, totals, accents, amounts, CSV, command options. The first time, the agent writes them; after that, the right move is to reuse what it wrote.
The criterion was fixed in writing before measuring: Sens leaves the project smaller only if all three sequences with Sens end below all three without it. With no real difference, that happens by chance one time in twenty.
Project size after each task
- Without Sens
- Reference
- With Sens
Lines of code in the project after each of the 30 tasks. Thin lines are each sequence; thick ones, their median. The dashed line is a reference solution written to reuse, which builds its shared modules early.
Show the data
| After task | Without Sens | Reference | With Sens |
|---|---|---|---|
| 0 | 88 | 88 | 88 |
| 1 | 92 | 93 | 93 |
| 2 | 102 | 113 | 103 |
| 3 | 113 | 150 | 109 |
| 4 | 121 | 161 | 118 |
| 5 | 145 | 188 | 137 |
| 6 | 162 | 202 | 161 |
| 7 | 178 | 209 | 177 |
| 8 | 223 | 234 | 195 |
| 9 | 232 | 243 | 206 |
| 10 | 232 | 243 | 206 |
| 11 | 258 | 276 | 231 |
| 12 | 277 | 282 | 247 |
| 13 | 296 | 307 | 265 |
| 14 | 304 | 309 | 269 |
| 15 | 316 | 315 | 271 |
| 16 | 346 | 345 | 301 |
| 17 | 354 | 357 | 309 |
| 18 | 367 | 369 | 326 |
| 19 | 370 | 377 | 330 |
| 20 | 381 | 390 | 341 |
| 21 | 398 | 404 | 357 |
| 22 | 415 | 413 | 373 |
| 23 | 426 | 422 | 384 |
| 24 | 432 | 422 | 388 |
| 25 | 452 | 434 | 401 |
| 26 | 453 | 436 | 402 |
| 27 | 476 | 453 | 418 |
| 28 | 485 | 455 | 426 |
| 29 | 487 | 455 | 427 |
| 30 | 496 | 455 | 436 |
Size at task 30, each sequence
All three sequences with Sens end below all three without it. The criterion holds: median 436 lines against 496, 12% smaller, with a 95% interval of −140 to −28 lines.
Tokens spent, accumulated over the 30 tasks
- Without Sens
- With Sens
With a smaller project to read at every task, Sens spends less: 27.7 million tokens across its three sequences against 33.7 million. Tokens include cache reads, so they measure the volume of work, not exact cost.
Show the data
| After task | Without Sens | With Sens |
|---|---|---|
| 0 | 0.0M | 0.0M |
| 1 | 0.3M | 0.3M |
| 2 | 0.5M | 0.6M |
| 3 | 0.9M | 0.8M |
| 4 | 1.2M | 1.1M |
| 5 | 1.6M | 1.3M |
| 6 | 1.8M | 1.7M |
| 7 | 2.0M | 1.9M |
| 8 | 2.5M | 2.4M |
| 9 | 3.4M | 2.7M |
| 10 | 3.7M | 3.0M |
| 11 | 4.4M | 3.4M |
| 12 | 4.7M | 3.6M |
| 13 | 5.2M | 4.0M |
| 14 | 5.4M | 4.2M |
| 15 | 5.9M | 4.5M |
| 16 | 6.2M | 5.1M |
| 17 | 6.6M | 5.3M |
| 18 | 6.9M | 5.7M |
| 19 | 7.1M | 6.1M |
| 20 | 7.5M | 6.4M |
| 21 | 8.1M | 6.7M |
| 22 | 8.5M | 7.1M |
| 23 | 8.8M | 7.4M |
| 24 | 9.0M | 7.7M |
| 25 | 9.4M | 7.9M |
| 26 | 9.8M | 8.2M |
| 27 | 10.2M | 8.5M |
| 28 | 10.7M | 8.8M |
| 29 | 11.0M | 9.1M |
| 30 | 11.3M | 9.3M |
Where the difference comes from
Not from copying less. Neither arm copied blocks in earnest: jscpd found 0, 6 and 6 duplicated lines without Sens and none with it, and the probes for each planted concept give the same or nearly the same counts in both. The difference comes from writing less for the same thing. With Sens, the agent writes more functions, and shorter ones: a median of 28 against 20. In the pilot, to read quoted descriptions in the CSV, the agent without Sens wrote a whole CSV reader, 99 lines; with Sens, it noticed the description was the last field and needed two one-line functions.
Total tokens: 33.7M without Sens · 27.7M with Sens
Single tasks
Twelve tasks in three languages on two real projects, Sens itself in TypeScript and Rust and the Python library click at fixed commits, each validated against hidden tests and a reference solution. Three conditions with the same model: Claude Code alone (C0), the Canon as text in the system prompt with no circuit (C1), and Sens whole (C2).
| Measure | C0 · alone | C1 · Canon as text | C2 · Sens |
|---|---|---|---|
| Valid runs | 32/36 | 31/36 | 67/72 |
| Runs that added tests | 23/36 | 36/36 | 72/72 |
| Reused plain, far from the edit | 0/3 | 3/3 | 6/6 |
| Reused titleOf, far from the edit | 1/3 | 1/3 | 6/6 |
| Solved py-progress-final | 0/3 | 0/3 | 3/6 |
The text alone already gets much of the reuse when the helper is near. It does not get the cases where the helper is far away and named differently, titleOf, 1 of 3 against 6 of 6, nor the task that needs the shared cause fixed instead of one path, py-progress-final, 0 of 3 against 3 of 6. In single tasks, lines of code are noise: C2 writes about two lines fewer per task, but the interval touches zero. That variability is why Horizonte exists.
Runs that added at least one line of test
With Canon 1.0 the agent nearly stopped writing tests: it read “do what was asked and nothing more” as forbidding them, and took Sens's approval for a test run. Canon 1.1 says both things that were missing: a test that proves the change is part of the change, and Sens's approval is not a test run. Claude Code alone does not get the Canon, so its bars are each batch's baseline.
The rules
The change rules are deterministic. They compare fingerprints of every function, method and class, and of every four statements in a row, built by a Rust index over tree-sitter that keeps the project in memory: Sens's own repository, 556 files and 10,000 units, indexes in under two seconds, and looking up the copies of a unit takes about a microsecond. Exact copies and copies with renamed names match by hash; copies with lines added or removed, by MinHash over normalised tokens. The 0.80 threshold and the 80-token floor for blocking come from editing 400 functions of a real repository and reviewing every match by hand.
| Rule | Catches | Answer |
|---|---|---|
| R1 Reuse | A new function, method or class with the same type 1 or type 2 fingerprint as an existing one | Blocks from 80 tokens; below that, Claude is asked to think again |
| R2 Near copy | Type 3 similarity over the threshold, or a small function that matches another in shape and vocabulary | As R1; a note in tests |
| R3 New dependency | A manifest gains a dependency, in ten formats | Asks you |
| R4 Orphans | A new symbol nothing reaches, or an existing one the turn left unused | Blocks if internal; a note if exported |
| R6 Project rules | Rules you declare; the first is no comments | Blocks |
| R7 Integrity | Writing to .sens/, .git/, .claude/settings*.json or .mcp.json, or git worktree | Always blocks; restored if it came through the shell |
| R8 Protected tests | The turn removes tests or assertions Sens had approved | Asks you |
Rules cannot see judgement errors. For those, when a turn that touched code passes the rules, a reviewer reads the diff with the candidates the index found. Its output is not trusted blindly: any finding whose quote is not literally in the diff is dropped, and only high confidence blocks.
| Note | Catches |
|---|---|
| S1 | An abstraction with no second use |
| S2 | A fix to the symptom instead of the cause |
| S3 | Reinventing what the platform or a dependency gives |
| S4 | Speculation: options or branches nobody asked for |
| S5 | Clever where plain was enough |
| S6 | A dangerous cut: validation, error handling or security removed |
| S7 | Reinventing what the project has, citing a candidate |
What did not work
Every block the circuit made was reviewed by hand, with its diff and its conversation. Blocks are rare, seven in 228 runs of Sens, so a single unfair one weighs a lot. We aimed for fewer than 5% unfair blocks and did not reach it in any batch that had blocks; each unfair one had a concrete cause, now fixed with a test that pins it.
| Batch | Blocks | Unfair | Cause | Fix |
|---|---|---|---|---|
| Hard tasks, C2 v3 | 1 | 0, 1 debatable | The reviewer flagged an idiom the project repeats | S7 on something private is only a note |
| Calibration, C2 | 2 | 2 | R8 compared with the file before each write | R8 compares with the last approved state |
| Calibration, C2 after the fix | 0 | 0 | — | — |
| Horizonte, pilot | 1 | 0 | — | — |
| Horizonte, confirmation | 3 | 1 | R8 took a test helper for a test | Only what checks something counts as a test |
A wrong list is worse than none. The first version of Sens suggested eight symbols unrelated to the request; the model read them, searched no further and rewrote the accent helper by hand three times, while the Canon as text, with no list, imported it three times. With the search rebuilt, the helper appears among the suggestions and Sens uses it every time, without blocking anything.
Limitations
- One model. Every run used Claude Sonnet 5.5 at medium effort.
- One project, one language, three sequences per arm. The confirmatory criterion is strict, all below all, but the size of the effect has a wide interval.
- We wrote the tasks. So they could not tilt the result, the tasks, their tests and the reference were committed before the first run, and the criterion was fixed before measuring.
- No MCP tools in the benchmark. Sens was measured without the index queries the app offers, so the result is a lower bound.
- Nineteen languages not yet benchmarked. Vue, Svelte and the languages added later are covered by tests, not by agent runs.
- Tokens include cache reads. They measure the volume of work, not exact cost.
- The reviewer is not precise. Of the seven notes and blocks of its own that we reviewed, five were wrong. Its notes stop nothing, but they reach the model and you.
- Unwritten conventions. Sens does not know a project's implicit rules, such as keeping heavy imports inside a function.
Method
456 agent runs in five batches: a pilot, three hard tasks on Sens itself, twelve calibration tasks, and Horizonte's pilot and confirmation. Every condition used claude-sonnet-5-5 at medium effort, with Claude Code in --safe-mode so the author's own configuration could not leak into the runs. Every task was validated before use: the project's tests pass and the hidden ones fail at the start, and both pass with the reference applied. A regression has to fail twice in a row to count. Differences are medians with a 95% bootstrap interval, 10,000 resamples with a fixed seed; Horizonte's criterion is an exact permutation test.
Reproduce it
sens-bench validate --tasks bench/tasks
sens-bench run --tasks bench/tasks --condition C0,C1,C2 --reps 3 --out bench/results/<batch>
sens-bench sequence validate bench/sequences/cuentas
sens-bench sequence run bench/sequences/cuentas --condition C0,C2 --reps 3 --out bench/results/<batch>
sens-bench sequence report bench/results/<batch>Each run's data, its diff and the tasks are in the bench/ folder of the Sens repository.
The Canon
The text every session receives, word for word, in English as the model reads it. The circuit is what makes it more than advice.
# Sens Canon v1.1
You are working inside Sens. Sens indexes this project and judges every change you make before your turn can end. What Sens tells you about this project, in its messages, denials and reviews, is a fact about the code, not a suggestion. When Sens names something to reuse, reuse it.
## Before you write
Go down this ladder and stop at the first step that answers the need:
1. Is it needed? Do what the person asked and nothing more: no speculative options, parameters, flags or branches. A test that proves the change is part of the change, not something extra.
2. Does the project already have it? Reuse the existing function, component, type or constant. Ask Sens with `already_exists` or `find_symbol` when unsure.
3. Does the standard library or the platform give it? Use that.
4. Does an installed dependency give it? Use that. A new dependency needs the person's approval, and Sens asks them for it.
5. Only then write new code: the smallest version that is correct.
## While you write
- Fix the cause in the shared code, not the symptom in each caller.
- No abstraction without a second real use: no interface, factory, wrapper, layer or configuration for a single consumer.
- Boring over clever. Match the names, patterns and style of the code around you.
- If you would copy a block, extract it once and call it from both places.
- Delete what your change leaves unused.
## Never cut
Less code never means removing validation at trust boundaries, error handling that prevents data loss, security checks, accessibility, or anything the person asked for.
It never means skipping tests either. When your change alters behaviour and the project has tests, add or extend one that fails without your change, in the style of the tests around it, and run the tests you touched before you finish.
## Working with Sens
- A denied write comes with the reason and what to use instead. Change the approach. Retrying the same thing through the shell, another tool or a subagent does not help: Sens judges what lands on disk, however it got there.
- When you finish, Sens audits the whole turn. If it blocks, fix what it found and finish again.
- Sens judges the shape of the code, not whether it works. Its approval is not a test run: that part is yours.
- Never edit `.sens/`, `.claude/settings*.json` or `.mcp.json`.