Forge
The ROI of Forge: Measuring Developer Velocity in Defense Tech
Why the return on a sealed coding appliance shows up in cleared-engineer throughput and avoided rework, not in a subscription line item, and how to build a measurement framework that survives an audit.
· 8 min read
Ask a program manager in defense tech to justify a coding-assistant purchase and they will reach for the same ROI model a commercial software team would use: hours saved per engineer, multiplied by loaded cost, compared against license cost. That model breaks down almost immediately against a cleared engineering team, not because the arithmetic is wrong but because every term in it means something different. Hours are not interchangeable when the pool of people allowed to spend them is fixed by clearance level rather than headcount budget. License cost is not the comparison point when the alternative to a sealed appliance is not "a cheaper tool" but "no AI-assisted coding at all, because nothing else clears security review." The return on Forge has to be modeled on its own terms, using the constraints that actually bind in a classified or export-controlled engineering organization.
The scarce resource isn't hours, it's cleared hours
In commercial software, developer time is fungible: if a team is under-resourced, it can hire, contract, or reassign someone from an adjacent project, usually within a quarter. In defense engineering, the binding constraint is not budget but clearance. A TS/SCI-cleared systems engineer cannot be conjured by increasing a line item, and the pipeline that produces one runs on a timeline measured in months to years, largely outside the program's control. That changes what "productivity" means. A velocity gain that lets a cleared team absorb work that would otherwise require hiring another cleared engineer is not a convenience. It is relief on the actual constraint the program is fighting, and it should be modeled against the cost and delay of adding headcount, not against an engineer's hourly rate.
This reframes the whole ROI conversation. A coding assistant that saves a commercial team hours is competing against the option of just hiring more people. A coding assistant that saves a cleared defense team hours is competing against an option that frequently does not exist on the program's timeline at all. The relevant baseline is not "engineer minus tool cost." It is "cleared engineer capacity that would otherwise require a req, a clearance adjudication, and a wait."
Where Forge changes the denominator, not just the numerator
Most coding-copilot ROI arguments focus on the numerator: faster autocomplete, fewer boilerplate hours, quicker onboarding to an unfamiliar codebase. Those effects are real and Forge produces them the way any capable code-aware assistant does — repository-indexed context tends to make suggestions more relevant than a context-free autocomplete, and that shows up as fewer discarded suggestions and less time spent explaining the codebase to a new hire. But for a cleared team, the more consequential effect is on the denominator: how much of an engineer's day is actually available for engineering work in the first place.
A cleared engineer working in an air-gapped enclave loses time to friction that a commercial developer never encounters: pulling reference material through an approved transfer process, working without the autocomplete and search tooling they use on personal projects because those tools assume internet access, context-switching between an isolated dev environment and whatever system holds documentation. None of that shows up as "time spent coding" in a naive productivity model, but it is real time lost to the boundary itself. A tool that lives inside the enclave and behaves like the assistant an engineer would use anywhere else, because it can, without needing an exception, recovers some of that boundary tax directly. That recovery is specific to sealed, air-gapped environments and does not show up at all in ROI models built for a team working on an ordinary internet-connected laptop.
The rejected-tool cost that never appears on a spreadsheet
The most consequential line item in a defense-tech ROI model is usually the one nobody puts on the spreadsheet: the cost of the coding assistants that never made it through security review at all. A cloud-based copilot that sends prompt context to a vendor endpoint is not a competitively priced alternative to Forge for a team handling ITAR-controlled or classified source. It is disqualified before evaluation, and the real comparison is not Forge-versus-cloud-tool but Forge-versus-nothing. Modeling ROI against a disqualified competitor understates the return substantially, because it implies engineers had a cheaper AI-assisted option and chose the sealed one at a premium, when in most defense programs no such option was ever on the table.
The corollary is that the accreditation timeline itself belongs in the ROI model, not just the per-engineer productivity delta. A tool that takes a security team eight months to approve, revisit, and finally reject represents eight months of program time spent on an evaluation that produced no usable capability. That's an opportunity cost a sealed appliance whose data-flow story is structural rather than contractual avoids by design. An architecture where the honest answer to "what leaves the enclave" is "nothing, because there is no connection out" tends to compress that review cycle, and the compression itself is a return worth counting, separate from anything an engineer experiences day to day.
What to actually instrument
A measurement framework worth defending to a program office needs to track a small number of things consistently, not produce one aggregate efficiency number that nobody can decompose later.
Suggestion acceptance rate, tracked over time and by task type, is the closest proxy available for whether the assistant is actually useful on this codebase rather than generically capable — a rate that climbs as the repository index matures indicates the tool is learning the codebase, which is a different and more durable signal than raw usage volume.
Time-to-first-commit for new team members, compared before and after Forge is available, captures the onboarding effect directly and tends to be one of the more legible numbers to present upward, since ramp time is already a metric most engineering programs track for other reasons.
Rework and defect-escape rate on AI-assisted commits versus unassisted ones, tracked through existing code review and defect data rather than a bespoke study, answers the question every skeptical reviewer eventually asks: is this actually producing better work, or just faster work that costs more to fix later. This is the metric most likely to reveal a genuine problem if one exists, and a program that only tracks speed and never checks this is measuring half the equation.
Boundary-friction time, the portion of an engineer's day spent on enclave-specific overhead rather than engineering, is the hardest of these to instrument well because it usually is not tracked at all today, which means establishing a pre-Forge baseline requires deliberate effort before rollout, not after. It is also the measurement most specific to this environment, and skipping it because it is inconvenient to collect throws away the number most likely to distinguish a sealed-appliance ROI story from a generic copilot one.
None of these should be reported as a single figure like "N percent faster." A program office evaluating a capital and operational decision deserves the decomposed picture — which effect is driving the return, and whether it is a one-time onboarding gain or a durable per-sprint one — because that is the version of the number that survives a follow-up question in a budget review.
Building the baseline before you need it
The single most common mistake in measuring Forge's return is starting the measurement after rollout, which leaves no honest before-state to compare against and forces the program to rely on engineer self-report months later, a notoriously unreliable instrument. The baseline — current time-to-first-commit, current rework rate, current boundary-friction time if it can be estimated at all — should be captured deliberately in the weeks before deployment, using the same instrumentation that will run afterward. This is not a large undertaking; it is a discipline, and it is the difference between an ROI claim a program can defend in an audit and one that rests on people's memory of how things used to feel.
It is also worth stating plainly what this framework does not do: it does not produce a number that transfers between programs. A team's clearance mix, codebase size, boundary friction, and baseline defect rate are all program-specific, and a velocity gain measured on one team's monorepo says little about what another team with a different profile should expect. Treat any generic industry figure — including any figure that sounds like it came from this document — with the same skepticism a rigorous engineer would apply to an unsourced benchmark, and build the program's own baseline instead of borrowing someone else's.
The honest shape of the return
The ROI case for Forge in a defense engineering program is not that it makes engineers write code faster in the way a general productivity tool might. It is that it recovers time currently lost to the specific friction of working behind an air gap, it converts a category of tooling that was previously disqualified from evaluation entirely into one that clears review, and it does both without asking the program to accept a data-flow risk in exchange for the gain. Whether that return is worth the appliance's cost is a question every program has to answer with its own numbers: the scarcity of its cleared engineers, the size of its boundary tax, the actual cost of the accreditation cycles it has burned on tools that never shipped. What a defensible ROI model owes the decision is not a bigger number. It is the decomposed, program-specific one, measured against a baseline captured before the tool arrived.