Measurement
The Three Velocities of Effective AI-Assisted Engineering
More output is useful only when the codebase remains changeable and the team can review, understand, and maintain what it ships.
The Three Velocities of Effective AI-Assisted Engineering
- 01FlowWork reaching completion
- 02StructureSafe future changes
- 03PracticesLearning and review
- 04EvidenceSources and limitations
Adapted for Binomial from Ravi Singh’s The Three Velocities of AI Coding: Quantitative, Structural, and Behavioral, originally published February 2, 2026.
A team introduces AI coding tools. Changes get drafted faster, pull requests arrive more frequently, and the backlog appears to move. Then review queues lengthen. Engineers find competing implementations of the same concept. The team has accelerated one part of delivery while making another harder.
Ravi Singh’s three-velocity framework gives leaders a way to examine that imbalance. For Binomial’s audience, its value is practical: look at delivery flow, codebase integrity, and team practices together. These are three perspectives on engineering effectiveness, rather than three numbers to combine into a productivity score.
Delivery flow: does useful work reach users?
Quantitative velocity describes the amount of work moving through the system. AI can shorten drafting, but drafting is only one interval. A change still needs review, validation, integration, and release. More open pull requests may mean more work waiting, rather than more value delivered.
Start with a defined window and comparable changes. Examine completed changes, time to first review, time waiting for review, and time to merge. If deployment records are available and in scope, follow the work through release. A merge timestamp alone cannot establish when a customer received a change.
Suppose small changes reach review sooner, but large changes wait longer. Investigate whether the team has moved effort from authors to reviewers. An appropriate intervention could be smaller change boundaries and clearer descriptions, followed by another observation window. Increasing generation capacity would leave the constraint unresolved.
Codebase integrity: can the next change stay safe?
Structural velocity concerns the codebase’s capacity to support continued change. A locally plausible implementation can still duplicate an existing abstraction, bypass an ownership boundary, or leave a migration incomplete. Those choices create work for the next engineer.
Inspect recurring follow-up fixes, duplicated patterns, dependency changes, test coverage of important behaviors, and the explanations behind architectural exceptions. Recovery time and change failure rates are useful where reliable incident and deployment records exist. Repository activity by itself does not supply either measure.
A practical review asks whether the change reuses established patterns, tests the failure path, and preserves a clear interface. A higher churn count is a prompt to inspect the diff and its context: intentional refactoring and repeated correction can produce similar activity patterns.
Team practices: can people use the tools well?
Behavioral velocity concerns learning, judgment, and coordination. Engineers need to define bounded tasks, supply useful context, review model output, and explain their decisions. Reviewers need enough time and domain knowledge to challenge the result.
Use engineering conversations alongside repository evidence. Ask where assistance reduced effort, where generated work required correction, and which practices others could reuse. Tool adoption and developer experience require their own sources; neither can be inferred reliably from a contributor’s commit count.
Coaching should address a specific practice. For example, a developer might benefit from writing acceptance checks before generation, while a reviewer might need clearer ownership of cross-service changes. Describe the behavior and its context without turning a diagnostic into an employee ranking.
Use the three perspectives in one review
Take a sample of changes and review the same sample through all three lenses. Record what happened, what you think explains it, and what additional evidence would distinguish competing explanations. Differences in change size, release pressure, and repository maturity belong in that discussion.
- Flow improved, integrity weakened: examine review depth, test gaps, and unnecessary additions.
- Integrity improved, flow stalled: check whether deliberate remediation or waiting time explains the delay.
- Tools are used, practices vary: share an effective workflow and observe whether it helps comparable work.
Choose one intervention, name its owner, and define the next review date. Keep the baseline and comparison scope visible. A useful outcome is a better operating decision, with enough evidence to revisit it.
How this informs a Binomial assessment
The AI Engineering Readiness Assessment connects individual coaching, team capability, repository readiness, and technical-debt priorities. This framework helps organize those questions around the constraints on safe delivery.
GitHub-observable evidence can support a review of change and review patterns. It does not independently identify which tool produced a change or prove that AI caused an improvement. Keep observed evidence, modeled explanations, confidence, and missing sources explicit before making an investment decision.
Key takeaways
What to put into practice
- Follow work beyond generation
- Review codebase integrity with flow
- Combine repository evidence with team context
Working framework
Questions for the next review
Original essay
Ravi Singh • February 2, 2026
This Binomial adaptation offers editorial guidance. It does not report measured Binomial customer results.