Field Notes
The Salary Slider Already Knows The Answer
A new Copilot dashboard places compensation beside pull-request output. The arithmetic is directional, but the management theory points one way.
An engineering leader opens a dashboard and chooses a salary band.
The cards recalculate. One shows developers who mainly use chat and code completion. The other shows developers working agent-first. Each card places Copilot cost beside a percentage of payroll and an average number of pull requests per developer per month.
The interface calls this potential return on investment.
GitHub is careful about the claim. Its new Copilot ROI section uses AI-credit consumption for cost, a selected compensation band rather than anyone's actual salary, and pull-request output as the return. The documentation says the figures are directional estimates, not financial results.
That caveat matters. So does the shape of the calculator around it.
The dashboard divides people into adoption phases. Passive and code-first users sit on one side. Agent-first and multi-agent users sit on the other. The call to action is explicit: compare the groups, justify the investment, and find the people with enough “headroom” to move deeper into adoption.
The salary slider can change the price of the labor. It cannot change the theory of the work.
In that theory, more pull requests are the visible return. A developer who uses more agentic tools and merges more pull requests appears farther along. The calculation does not say that person is better, and it should not be read as an individual performance score. Still, an interface does not have to make a formal claim in order to teach a direction. One cohort is named as a phase to leave. Another is the phase to reach.
I have spent enough time inside software teams to know how quickly a directional metric acquires a destination.
A developer can spend a week deleting code, preventing a migration, helping a teammate understand an old subsystem, or discovering that the requested feature should not exist. The pull-request count may be one, zero, or negative in spirit. Another developer can divide a broad change into six clean branches and merge all six. The second week produces more visible units. The first may have saved the organization from owning a seventh.
Neither example proves that pull-request counts are useless. They are useful records of motion. They can help a team see whether work is getting stuck, whether review has become a bottleneck, or whether an agent-assisted workflow is changing the size and cadence of changes. Trouble begins when a record of motion is asked to stand in for the value of arriving somewhere.
This is the more literal sequel to AI Metrics Are Management Design. That note argued that usage dashboards become small theories of what an organization values. The salary selector makes the theory unusually easy to inspect. Cost is denominated in dollars. Return is denominated in pull requests. The unpriced remainder includes maintenance avoided, incidents prevented, product judgment, customer consequence, learning, and the review work transferred to somebody else.
GitHub is not ignoring quality. One of its newest releases adds organization-level code quality trends: open findings over time, grouped by health score or severity, with tables showing which repositories are improving and which may need support. The underlying system combines deterministic CodeQL rules, coverage checks, and AI analysis of recently changed code.
That is a more useful counterweight than pretending every pull request is equal. It still does not complete the equation. A falling count of recognized findings can coexist with the wrong product, an exhausted team, a brittle architecture, or a beautifully maintained service nobody needed. Quality is wider than the dashboard too.
Software engineering has had this argument for years. The SPACE framework treats developer productivity as a multidimensional system involving satisfaction, performance, activity, communication, and flow. Its durable warning is that no single measure can carry the whole idea. DORA's research on AI-assisted software development reaches the problem from another side: AI tends to amplify the conditions already present in an organization. A strong feedback system may become stronger. A confused delivery system can produce confusion faster.
The current generation of enterprise AI reports keeps rediscovering the appeal of an imperfect proxy. OpenAI's new Enterprise Signals ranks “frontier” firms partly by output tokens per active user. The report plainly says token volume is an imperfect measure of business value—a short response may matter more than a long one—then uses the proxy to describe a widening gap between frontier and typical firms.
The disclaimer is honest. The leaderboard remains legible.
This is how measurement becomes office weather. A cautious proxy enters a report. The report names a leading cohort. The cohort becomes a maturity model. A manager receives a benchmark and a recommendation to close the gap. Soon restraint has to explain itself, while greater use arrives pre-approved as progress.
There is a better way to use these instruments. Keep the operational signal, but make the organizational question harder to skip. When pull-request output rises, ask what changed for the people reviewing it. When agent cost rises, ask which work became possible and which obligation should disappear. When a team moves into a deeper adoption phase, look for a product outcome, a maintenance gain, or a calmer delivery system—not just more artifacts crossing a merge boundary.
The right comparison may also reveal that one team should use less. Their work may be unusually sensitive, difficult to verify, or already constrained by review capacity. Maturity should include knowing when another agent run would create more residue than leverage.
An ROI model is allowed to be partial. Every model is. The honest design task is to keep the missing parts close enough that nobody mistakes a clean calculation for a complete bargain.
The salary slider can estimate what the tool costs beside human labor. It cannot decide what the human labor was for. If the dashboard already knows which kind of developer should win before the work is examined, the answer is not return on investment. It is the organization's preference, rendered as arithmetic.