Ask a software engineer whether AI made them faster this year and 84% will say yes. Ask their VP of engineering to put a number in front of the board and 39% have nothing to show. Both figures come from the same study: GitKraken’s State of AI in Engineering 2026, published on August 20 and built on responses from 554 developers and engineering leaders.
That gap lands harder in SaaS than in most industries, because research and development is the largest line in the operating model. SaaS Capital’s 2026 spending benchmarks, drawn from a March 2026 survey of more than 1,000 private B2B SaaS companies, put median R&D spend at 22% of ARR. Selling costs came in at 15%. Marketing at 8%. The biggest cost center in a SaaS company carries the thinnest reporting, which is why developer experience metrics are appearing in board packets next to net revenue retention and CAC payback.
The largest line in the SaaS model is the least instrumented
Sales efficiency has three decades of convention behind it. A board can ask for CAC payback by segment, pipeline coverage against next quarter’s number, or win rate by ACV band, and get answers a new investor reads without a tutorial. Ask the same board for the return on the R&D line and the answer is a headcount plan and a roadmap.
The proportions make that awkward. In the SaaS Capital data, R&D at 22% of ARR is the single largest departmental median, ahead of selling costs and general and administrative expense at 15% each. Smaller companies carry more of it: in the $3 million to $5 million ARR band, R&D runs at 24% of revenue. Equity-backed companies spend 56% more on R&D than bootstrapped peers, precisely the cohort whose boards meet quarterly.

Public filings make the point in dollars. Atlassian spent $2.67 billion on research and development in the fiscal year ended June 30, 2025, against $5.22 billion of revenue, per its fourth quarter fiscal 2025 results filed with the SEC. That is 51% of revenue on engineering against 22% on marketing and sales. In fiscal 2024 the ratios were 50% and 20%.
No serious operator wants that line smaller. Product velocity is the asset. The argument is narrower: a company that can explain what a dollar of sales spend bought, and cannot explain what a dollar of engineering spend bought, is running its biggest bet on faith.
What the proof gap actually looks like
The GitKraken numbers describe an industry that finished the adoption question and never started the measurement one. Tool adoption reached 96.4% of teams. Sentiment is equally settled: 84% of developers say AI made them more productive, 43% say much more productive, and fewer than 5% report feeling slower.
Then the reporting collapses. Only 20% of organizations measure productivity in any specific way. Thirty-nine percent have no way to measure AI’s impact at all, and another 33% rely entirely on developers self-reporting that it helped. That is 72% of engineering organizations “running on belief instead of a number”, in the researchers’ phrasing.
Scale grades the problem. As organizations grow, the share with no way to measure AI’s impact falls from 47% to 19%, and the share using DORA or comparable engineering metrics climbs from 6% to 21%. Four in five large engineering organizations still lack a delivery metric set, and the work is outrunning the reporting: developers naming agent delegation as their primary way of working went from 7.6% in September 2025 to 28% in June 2026.
The financial framing is usually wrong, and it matters. Tomasz Tunguz’s June 2026 analysis of AI spend per engineer put the median software company at roughly $137 per employee per year, against about $89,000 at the top 1% of firms. The tooling invoice is a rounding error. The number under scrutiny is the payroll it is supposed to make more effective.
Self-reported speed is the weakest evidence a board can accept
The strongest evidence here comes from a group that keeps publishing against its own prior results. In a randomized controlled trial run between February and June 2025, METR found that 16 experienced open source developers took 19% longer to complete 246 real issues when AI tools were allowed, with a confidence interval from 2% to 39%. Those developers had forecast a 24% speedup. After finishing, they still believed AI had made them 20% faster.
METR then did the honest thing. Its February 2026 update covered 57 developers, 143 repositories and more than 800 tasks, and said the new data “gives us an unreliable signal.” The point estimates had flipped, at 18% faster for returning participants and 4% faster for new recruits, but both confidence intervals straddle zero. The cause is selection: 30% to 50% of developers said they were skipping tasks they did not want to attempt without AI.
That is not an argument against AI-assisted development. It is evidence that perception and measurement have separated, and that self-reported time savings now carry almost no information. Atlassian’s 2025 State of Developer Experience report, based on 3,500 developers and managers across six countries, found 99% reporting time savings from AI and 68% saving more than 10 hours a week. The same respondents said 50% lose 10 or more hours a week to non-coding work. Developers spend 16% of their time coding.

The alignment number should worry a board most. In that same research, 63% of developers say leaders do not understand their pain points, up sharply from 44% the year before. Atlassian’s reading: leaders banked the AI time savings without fixing the friction underneath, so the savings went somewhere invisible.
The metric set boards are converging on
The frameworks have stopped competing, which is the useful news. DORA’s four delivery measures, deployment frequency, lead time for changes, change failure rate and failed deployment recovery time, remain the closest thing engineering has to GAAP. SPACE added what throughput counters miss, including satisfaction and collaboration. DX’s Core 4 folds all three into four counterbalanced dimensions: speed, effectiveness, quality and business impact. DX says the approach runs at more than 300 companies and reports a 3% to 12% gain in engineering efficiency. Treat that as a vendor-reported outcome; the structure is what travels.
Counterbalancing is what survives contact with a board. A speed metric alone invites gaming, so it travels with a quality metric. DX flags one of its own components, diffs per engineer, as requiring caution, and that caution has aged well: when agents author a share of commits, throughput counters inflate without value being created. GitHub’s Octoverse 2025 recorded nearly one billion commits in 2025, up 25.1% year over year. Commit volume has never been less informative.
Platform quality is the lever, not tool procurement
The most actionable finding of 2026 is not about tools. Google Cloud’s 2025 DORA report, based on roughly 5,000 technology professionals, found 90% use AI at work and more than 80% believe it raised their productivity. It also found 90% of organizations have adopted at least one internal platform, and a direct correlation between that platform’s quality and the ability to get value from AI. DORA’s summary of the year: AI acts as an amplifier, magnifying an organization’s existing strengths and weaknesses.
That maps onto the friction data. Atlassian’s respondents named finding information across services, docs and APIs, adapting to new technology, and context switching between tools as their top time wasters. Coding did not make the list in either 2024 or 2025. A coding assistant makes the 16% faster. The platform and documentation surface decide the other 84%, which is the argument behind platform engineering becoming one of SaaS’s fastest-growing categories.

The quality tail nobody has priced yet
Throughput without a stability counterweight produces a bill that arrives late. DORA’s 2025 data found a positive relationship between AI adoption and both throughput and product performance, reversing the prior year, while the negative relationship with delivery stability persisted. Trust tracks that experience. Stack Overflow’s 2025 Developer Survey found more developers actively distrust the accuracy of AI tools, at 46%, than trust it, at 33%. Sixty-six percent named solutions that are “almost right, but not quite” a top frustration, and 45.2% said debugging AI-generated code takes longer.
Code review data points the same way. CodeRabbit’s analysis of 470 open source GitHub pull requests found AI co-authored pull requests carried roughly 1.7 times more issues than human-only ones, with logic and correctness problems 75% more common and security vulnerabilities up to 2.74 times higher. TechCrunch reported in May 2026, citing The Information, that Uber spent its entire 2026 AI budget in four months with no measurable productivity increase.
Two caveats belong here. None of this argues for slowing adoption; it argues for pairing every speed metric with a defect and stability metric before the next planning cycle. And the savings do not stay in R&D. Agentic workflows move cost out of payroll and into inference, which lands in cost of revenue, so an improving R&D ratio can sit alongside the gross margin compression that AI COGS is creating across SaaS.
What to put on the next four board decks
The practical version fits on one slide. Report DORA’s four delivery measures as a trend, never as a quartile ranking against strangers. Pair them with a developer experience index from a recurring survey, so the friction data carries a number. Put one quality measure, change failure rate or defects per release, beside the throughput number rather than on a later page. Then report R&D as a percentage of revenue next to gross margin and revenue per employee, so the three check each other.
Two operator caveats keep this from going wrong. Below roughly 30 engineers, quartile benchmarking is statistical noise; a quarterly developer experience survey plus a single lead time trend tells a founder more. And no per-engineer measure should cross into a performance review. The moment diffs per engineer touches compensation, the number stops describing the system and starts describing what people believe it rewards. Developer experience metrics diagnose an organization; they do not score individuals.
Adoption is finished as a story. Every SaaS company has the tools and almost every developer likes them. The differentiated position over the next several quarters belongs to companies that can say, in a format a board and a diligence team both accept, what the 22% line produced.
Frequently asked questions
What are developer experience metrics, and how do they differ from DORA metrics?
DORA metrics measure software delivery: deployment frequency, lead time for changes, change failure rate and failed deployment recovery time. Developer experience metrics measure the conditions producing that delivery, including friction, flow, tooling quality and time lost to non-coding work. DX’s Core 4 combines both, pairing speed and effectiveness measures with quality and a developer experience index. The distinction matters because delivery metrics tell a board what happened, while experience metrics tell it why, and only the second points at a fix.
Should a SaaS board track developer productivity or developer experience?
Both, reported together, because either alone invites gaming. A pure productivity view rewards volume, and volume means little now that agents author commits. A pure experience view can look healthy while delivery stalls. The workable pattern is a counterbalanced set: one speed measure, one quality measure, one experience index, and the R&D ratio beside gross margin. Atlassian’s 2025 research found 63% of developers believe leaders do not understand their pain points, which is what a productivity-only view tends to produce.
How do you measure the return on AI coding tools?
Not through self-reported time savings, which the 2025 and 2026 research has effectively disqualified. METR’s randomized trial found experienced developers took 19% longer with AI while believing they were 20% faster, and METR’s own February 2026 follow-up called its newer estimates unreliable because of selection effects. A defensible approach measures delivery outcomes before and after adoption, tracks change failure rate alongside throughput, and reports the tooling cost against the engineering payroll it is meant to improve.
What is a healthy R&D spend as a percentage of revenue for a SaaS company?
SaaS Capital’s 2026 survey of more than 1,000 private B2B SaaS companies puts the median at 22% of ARR, unchanged from the prior year and rising to 24% in the $3 million to $5 million ARR band. Equity-backed companies spend 56% more than bootstrapped peers. Public companies vary widely by model, with Atlassian running R&D at 51% of revenue in fiscal 2025. The ratio matters less than whether it trends with output, gross margin and retention.
Do DORA metrics still work when AI agents write the code?
The delivery measures hold up better than the output counters. Deployment frequency, lead time, change failure rate and recovery time describe a system’s ability to ship safely, and that stays meaningful regardless of who or what wrote the change. Counts of commits, pull requests or diffs per engineer degrade quickly, because agentic workflows inflate them without adding value. DORA’s 2025 data also found the negative relationship between AI adoption and delivery stability persisting, which makes the stability half of the set more useful.
The bottom line
The 2026 evidence is unusually consistent. Adoption is universal, sentiment is positive, and the reporting underneath is close to absent, with 72% of engineering organizations running on belief rather than a number. R&D meanwhile sits at 22% of ARR for the median private SaaS company. The companies closing that gap are not buying more tools. They are pairing a delivery metric with a quality metric and an experience index, fixing the platform friction that eats most of a developer’s week, and reporting the R&D ratio next to gross margin so nobody mistakes a cost transfer for an efficiency gain.







