Research ·

Less code for the same work.

Coding agents write more than they need to: they rewrite what a project already has, and every turn adds a little. Sens puts a circuit around each turn of Claude Code that the model cannot switch off. Over 30 chained tasks, the project ended 12% smaller for the same features, with 18% fewer tokens.

12%
less code for the same 30 features
18%
fewer tokens over the whole sequence
90/90
tasks accepted, with Sens and without
0
tests of earlier tasks broken, in either arm

Claude Sonnet 5.5 at medium effort · three sequences per arm · criterion fixed before measuring

The problem

An agent working on a real project does not see the whole project. It sees what it searches for, and it searches for what occurs to it. When the helper it needs lives in another module under another name, it writes it again. When a task allows a short solution and a long one, nothing pushes it toward the short one.

The usual answer is written rules: a CLAUDE.md, a skill, the system prompt. They help, but they are advice. The model may not load them, may forget them halfway through a turn, or may decide they do not apply. We cannot touch the model's weights, but we can control three things: what the model sees, what it is allowed to do, and when it is allowed to finish. Sens uses all three at once.

The circuit

Claude Code accepts hooks as callbacks on its own stream protocol. Sens answers five of them inside its process, with one invariant: a change is born unapproved. Only an audit that passes, or a person, approves it.

  1. 01

    You send a message

    UserPromptSubmit

    Sens takes a checkpoint of the project and hands Claude up to eight things that already exist and relate to the request, each with its signature, file:line and number of uses.

  2. 02

    Before every write

    PreToolUse · Write, Edit

    Sens rebuilds the resulting file in memory and applies the change rules. It can deny the write before it happens, with the reason and the code to reuse.

  3. 03

    Before every command

    PreToolUse · Bash, PowerShell

    Sens protects its own paths and the configuration, and treats git commit and git push as the end of a turn: nothing unapproved gets committed.

  4. 04

    After every tool

    PostToolUse

    Whatever the tool, Sens looks for the files that changed on disk and applies the same rules to them.

  5. 05

    When the turn ends

    Stop · SubagentStop

    Sens audits the whole diff since the last approved point: change rules, orphaned code and, if code changed, the reviewer. With blocking findings Claude keeps working; after three rounds the turn is held for you.

Nothing to go around

Before building anything, we checked live that the mechanism holds against a model trying to slip past it: a disableAllHooks written mid-session does not stop the callbacks, subagents go through them, and a denied git commit leaves the repository as it was. Every shortcut has something that closes it:

ShortcutWhat closes it
Writing through the shell, Python or another MCP serverThe check of the disk after every tool, and the audit when the turn ends
Subagents and background tasksThey go through the same hooks; the turn is not approved while they run
Turning the hooks off or editing the configurationThe hooks live in Sens's process; rule R7 blocks and restores those files
Opening a worktree of its ownWorktree tools are denied, and R7 covers git worktree
Declaring the work finishedThe audit decides when a turn ends, not the model
Starting a new turn to escapeThe audit runs from the last approved point, so whatever is pending carries over
Committing unapproved workA commit is treated as the end of a turn
Looping foreverThree rounds, then the turn is held

The limit is explicit: the circuit closes the mistakes and shortcuts of a model, not those of a hostile program running on the same machine.

Results

Single tasks start from a clean project, and the harm of an agent that does not reuse is not in one task but in the sum. Horizonte measures the sum: a TypeScript command line for household expenses that starts at 88 lines, and 30 product requests in a fixed order, each building on the last. Eight concepts are planted that several tasks need without saying so: dates, months and weeks, totals, accents, amounts, CSV, command options. The first time, the agent writes them; after that, the right move is to reuse what it wrote.

The criterion was fixed in writing before measuring: Sens leaves the project smaller only if all three sequences with Sens end below all three without it. With no real difference, that happens by chance one time in twenty.

Project size after each task

  • Without Sens
  • Reference
  • With Sens
0200400600051015202530TaskWithout Sens · 496Reference · 455With Sens · 436

Lines of code in the project after each of the 30 tasks. Thin lines are each sequence; thick ones, their median. The dashed line is a reference solution written to reuse, which builds its shared modules early.

Show the data
After taskWithout SensReferenceWith Sens
0888888
1929393
2102113103
3113150109
4121161118
5145188137
6162202161
7178209177
8223234195
9232243206
10232243206
11258276231
12277282247
13296307265
14304309269
15316315271
16346345301
17354357309
18367369326
19370377330
20381390341
21398404357
22415413373
23426422384
24432422388
25452434401
26453436402
27476453418
28485455426
29487455427
30496455436

Size at task 30, each sequence

400440480520560reference · 455Without SensWithout Sens: 496 lines496Without Sens: 564 lines564Without Sens: 484 lines484With SensWith Sens: 436 lines436With Sens: 456 lines456With Sens: 424 lines424

All three sequences with Sens end below all three without it. The criterion holds: median 436 lines against 496, 12% smaller, with a 95% interval of −140 to −28 lines.

Tokens spent, accumulated over the 30 tasks

  • Without Sens
  • With Sens
0.0M2.5M5.0M7.5M10.0M12.5M051015202530TaskWithout Sens · 11.3MWith Sens · 9.3M

With a smaller project to read at every task, Sens spends less: 27.7 million tokens across its three sequences against 33.7 million. Tokens include cache reads, so they measure the volume of work, not exact cost.

Show the data
After taskWithout SensWith Sens
00.0M0.0M
10.3M0.3M
20.5M0.6M
30.9M0.8M
41.2M1.1M
51.6M1.3M
61.8M1.7M
72.0M1.9M
82.5M2.4M
93.4M2.7M
103.7M3.0M
114.4M3.4M
124.7M3.6M
135.2M4.0M
145.4M4.2M
155.9M4.5M
166.2M5.1M
176.6M5.3M
186.9M5.7M
197.1M6.1M
207.5M6.4M
218.1M6.7M
228.5M7.1M
238.8M7.4M
249.0M7.7M
259.4M7.9M
269.8M8.2M
2710.2M8.5M
2810.7M8.8M
2911.0M9.1M
3011.3M9.3M

Where the difference comes from

Not from copying less. Neither arm copied blocks in earnest: jscpd found 0, 6 and 6 duplicated lines without Sens and none with it, and the probes for each planted concept give the same or nearly the same counts in both. The difference comes from writing less for the same thing. With Sens, the agent writes more functions, and shorter ones: a median of 28 against 20. In the pilot, to read quoted descriptions in the CSV, the agent without Sens wrote a whole CSV reader, 99 lines; with Sens, it noticed the description was the last field and needed two one-line functions.

Total tokens: 33.7M without Sens · 27.7M with Sens

Single tasks

Twelve tasks in three languages on two real projects, Sens itself in TypeScript and Rust and the Python library click at fixed commits, each validated against hidden tests and a reference solution. Three conditions with the same model: Claude Code alone (C0), the Canon as text in the system prompt with no circuit (C1), and Sens whole (C2).

MeasureC0 · aloneC1 · Canon as textC2 · Sens
Valid runs32/3631/3667/72
Runs that added tests23/3636/3672/72
Reused plain, far from the edit0/33/36/6
Reused titleOf, far from the edit1/31/36/6
Solved py-progress-final0/30/33/6

The text alone already gets much of the reuse when the helper is near. It does not get the cases where the helper is far away and named differently, titleOf, 1 of 3 against 6 of 6, nor the task that needs the shared cause fixed instead of one path, py-progress-final, 0 of 3 against 3 of 6. In single tasks, lines of code are noise: C2 writes about two lines fewer per task, but the interval touches zero. That variability is why Horizonte exists.

Runs that added at least one line of test

Canon 1.0 · pilot and hard tasks

Without Sens (C0)14/18
The Canon as text (C1)9/18
Sens (C2)2/18

Canon 1.1 · calibration

Without Sens (C0)23/36
The Canon as text (C1)36/36
Sens (C2)36/36

With Canon 1.0 the agent nearly stopped writing tests: it read “do what was asked and nothing more” as forbidding them, and took Sens's approval for a test run. Canon 1.1 says both things that were missing: a test that proves the change is part of the change, and Sens's approval is not a test run. Claude Code alone does not get the Canon, so its bars are each batch's baseline.

The rules

The change rules are deterministic. They compare fingerprints of every function, method and class, and of every four statements in a row, built by a Rust index over tree-sitter that keeps the project in memory: Sens's own repository, 556 files and 10,000 units, indexes in under two seconds, and looking up the copies of a unit takes about a microsecond. Exact copies and copies with renamed names match by hash; copies with lines added or removed, by MinHash over normalised tokens. The 0.80 threshold and the 80-token floor for blocking come from editing 400 functions of a real repository and reviewing every match by hand.

RuleCatchesAnswer
R1 ReuseA new function, method or class with the same type 1 or type 2 fingerprint as an existing oneBlocks from 80 tokens; below that, Claude is asked to think again
R2 Near copyType 3 similarity over the threshold, or a small function that matches another in shape and vocabularyAs R1; a note in tests
R3 New dependencyA manifest gains a dependency, in ten formatsAsks you
R4 OrphansA new symbol nothing reaches, or an existing one the turn left unusedBlocks if internal; a note if exported
R6 Project rulesRules you declare; the first is no commentsBlocks
R7 IntegrityWriting to .sens/, .git/, .claude/settings*.json or .mcp.json, or git worktreeAlways blocks; restored if it came through the shell
R8 Protected testsThe turn removes tests or assertions Sens had approvedAsks you

Rules cannot see judgement errors. For those, when a turn that touched code passes the rules, a reviewer reads the diff with the candidates the index found. Its output is not trusted blindly: any finding whose quote is not literally in the diff is dropped, and only high confidence blocks.

NoteCatches
S1An abstraction with no second use
S2A fix to the symptom instead of the cause
S3Reinventing what the platform or a dependency gives
S4Speculation: options or branches nobody asked for
S5Clever where plain was enough
S6A dangerous cut: validation, error handling or security removed
S7Reinventing what the project has, citing a candidate

What did not work

Every block the circuit made was reviewed by hand, with its diff and its conversation. Blocks are rare, seven in 228 runs of Sens, so a single unfair one weighs a lot. We aimed for fewer than 5% unfair blocks and did not reach it in any batch that had blocks; each unfair one had a concrete cause, now fixed with a test that pins it.

BatchBlocksUnfairCauseFix
Hard tasks, C2 v310, 1 debatableThe reviewer flagged an idiom the project repeatsS7 on something private is only a note
Calibration, C222R8 compared with the file before each writeR8 compares with the last approved state
Calibration, C2 after the fix00——
Horizonte, pilot10——
Horizonte, confirmation31R8 took a test helper for a testOnly what checks something counts as a test

A wrong list is worse than none. The first version of Sens suggested eight symbols unrelated to the request; the model read them, searched no further and rewrote the accent helper by hand three times, while the Canon as text, with no list, imported it three times. With the search rebuilt, the helper appears among the suggestions and Sens uses it every time, without blocking anything.

Limitations

  • One model. Every run used Claude Sonnet 5.5 at medium effort.
  • One project, one language, three sequences per arm. The confirmatory criterion is strict, all below all, but the size of the effect has a wide interval.
  • We wrote the tasks. So they could not tilt the result, the tasks, their tests and the reference were committed before the first run, and the criterion was fixed before measuring.
  • No MCP tools in the benchmark. Sens was measured without the index queries the app offers, so the result is a lower bound.
  • Nineteen languages not yet benchmarked. Vue, Svelte and the languages added later are covered by tests, not by agent runs.
  • Tokens include cache reads. They measure the volume of work, not exact cost.
  • The reviewer is not precise. Of the seven notes and blocks of its own that we reviewed, five were wrong. Its notes stop nothing, but they reach the model and you.
  • Unwritten conventions. Sens does not know a project's implicit rules, such as keeping heavy imports inside a function.

Method

456 agent runs in five batches: a pilot, three hard tasks on Sens itself, twelve calibration tasks, and Horizonte's pilot and confirmation. Every condition used claude-sonnet-5-5 at medium effort, with Claude Code in --safe-mode so the author's own configuration could not leak into the runs. Every task was validated before use: the project's tests pass and the hidden ones fail at the start, and both pass with the reference applied. A regression has to fail twice in a row to count. Differences are medians with a 95% bootstrap interval, 10,000 resamples with a fixed seed; Horizonte's criterion is an exact permutation test.

Reproduce it

sens-bench validate --tasks bench/tasks
sens-bench run --tasks bench/tasks --condition C0,C1,C2 --reps 3 --out bench/results/<batch>
sens-bench sequence validate bench/sequences/cuentas
sens-bench sequence run bench/sequences/cuentas --condition C0,C2 --reps 3 --out bench/results/<batch>
sens-bench sequence report bench/results/<batch>

Each run's data, its diff and the tasks are in the bench/ folder of the Sens repository.

The Canon

The text every session receives, word for word, in English as the model reads it. The circuit is what makes it more than advice.

# Sens Canon v1.1

You are working inside Sens. Sens indexes this project and judges every change you make before your turn can end. What Sens tells you about this project, in its messages, denials and reviews, is a fact about the code, not a suggestion. When Sens names something to reuse, reuse it.

## Before you write

Go down this ladder and stop at the first step that answers the need:

1. Is it needed? Do what the person asked and nothing more: no speculative options, parameters, flags or branches. A test that proves the change is part of the change, not something extra.
2. Does the project already have it? Reuse the existing function, component, type or constant. Ask Sens with `already_exists` or `find_symbol` when unsure.
3. Does the standard library or the platform give it? Use that.
4. Does an installed dependency give it? Use that. A new dependency needs the person's approval, and Sens asks them for it.
5. Only then write new code: the smallest version that is correct.

## While you write

- Fix the cause in the shared code, not the symptom in each caller.
- No abstraction without a second real use: no interface, factory, wrapper, layer or configuration for a single consumer.
- Boring over clever. Match the names, patterns and style of the code around you.
- If you would copy a block, extract it once and call it from both places.
- Delete what your change leaves unused.

## Never cut

Less code never means removing validation at trust boundaries, error handling that prevents data loss, security checks, accessibility, or anything the person asked for.

It never means skipping tests either. When your change alters behaviour and the project has tests, add or extend one that fails without your change, in the style of the tests around it, and run the tests you touched before you finish.

## Working with Sens

- A denied write comes with the reason and what to use instead. Change the approach. Retrying the same thing through the shell, another tool or a subagent does not help: Sens judges what lands on disk, however it got there.
- When you finish, Sens audits the whole turn. If it blocks, fix what it found and finish again.
- Sens judges the shape of the code, not whether it works. Its approval is not a test run: that part is yours.
- Never edit `.sens/`, `.claude/settings*.json` or `.mcp.json`.