Skip to main content
Service

AI-Native Developer Experience

From under 10% test coverage to 40% with no QA team. From three-month onboarding to time-to-first-commit on day one.

The fastest way to accelerate an engineering organization is not to hire more engineers - it is to remove the friction that makes existing engineers slow. AI-Native DevEx is the discipline of operationalizing agentic AI as production infrastructure across the SDLC, building the internal developer platform (IDP) that supports it, and tracking the velocity metrics that prove it.

I run a multi-model strategy across Claude Code, GitHub Copilot, AWS Kiro, OpenAI Codex, Gemini, Bolt, Lovable, and Snowflake Cortex - all wired as first-class GitHub Actions pipeline stages for code generation, test synthesis, cloud architecture scaffolding, and documentation refresh. Multi-model fluency tracked as a leading indicator alongside delivery and quality metrics, so teams develop proficiency across providers with no single-vendor dependency. I have partnered with CPOs on Lovable and jointly with product teams on Bolt to push AI tooling beyond engineering into rapid product prototyping.

Measurable outcomes: 5x deploy frequency, 23% PR throughput gain, code-to-release cycle time down 40%, MTTR cut 30%, and new-engineer onboarding cut 70%.

Hire Me
AI-native developer experience and SDLC
Capabilities

What I Deliver

AI-Augmented Pipelines

Claude Code, GitHub Copilot, and AWS Kiro running as first-class GitHub Actions stages - code generation, test synthesis, cloud architecture scaffolding, and documentation refresh embedded in every CI/CD run.

Ship Fast, Ship Quality

Hitting a deadline with software that does not work is not shipping - it is creating a cleanup project for next quarter. AI-generated unit tests, Playwright UI regression suites, AI-assisted PR review, and SonarQube quality gates in CI/CD lifted coverage from under 10% to 40% with no dedicated QA team. Speed and quality, not one or the other.

Onboarding That Ships Day One

AI-assisted documentation refresh plus a lead-mentor program cut new-engineer onboarding 70% - measurable DevEx and time-to-first-commit improvement on every new hire.

Velocity Metrics That Matter

MTTR, deploy frequency, PR throughput, SLA attainment, onboarding time-to-productivity, and AI-assisted deployment rate - tracked, trended, and tied to engineering investments. Board-ready, not vague AI-adoption slides.

Multi-Model Fluency

No single-vendor dependency. Teams build proficiency across Claude, GitHub Copilot, AWS Kiro, OpenAI Codex, Gemini, Bolt, Lovable, and Snowflake Cortex. Prompt throughput tracked as a leading indicator alongside delivery metrics.

IDP Build-Outs

Internal developer platforms with self-service CI/CD, CloudWatch Synthetics canaries, PagerDuty, custom Grafana dashboards, secrets management, and MCP servers giving AI agents governed access to the same tools engineers use.

Internal developer platform with self-service CI/CD and observability

Internal Developer Platform

The IDP is the foundation of AI-native DevEx. I build platforms that give engineers self-service access to everything they need - eliminating the operational friction that slows teams down and shifting the org from reactive firefighting to proactive observability.

  • Self-service CI/CD with AI stages: Engineers trigger deployments and create environments without infrastructure tickets. Claude Code, GitHub Copilot, and Kiro run as first-class pipeline stages on every build.
  • Observability as standard: CloudWatch Synthetics canaries, PagerDuty, Datadog, Splunk, and custom Grafana dashboards built into every service by default - executive-level visibility into uptime and platform health for the first time.
  • Secrets management: Vault-backed secrets accessible to services and AI agents through governed access controls.
  • Agent tooling integration: MCP servers embedded in the IDP so AI agents have governed, auditable access to the same tools engineers use.
AI Tooling

Multi-Model AI-Augmented Toolchain

The best AI toolchain is one that engineers actually use. I implement production-grade tooling that integrates with existing workflows rather than requiring engineers to change how they work - and I run it as a multi-model strategy so the org develops fluency across providers.

  • Claude Code, GitHub Copilot & AWS Kiro

    Frontier coding agents configured with codebase context, company-specific rules, and MCP server access - so every engineer has an AI pair programmer that knows the stack. Early-adopter design partner for AWS Kiro before its general availability.

  • OpenAI Codex, Gemini, Bolt, Lovable & Snowflake Cortex

    OpenAI Codex for prototyping and code review, Gemini for long-context reasoning, Bolt and Lovable for rapid app scaffolding alongside CPOs and product teams, Snowflake Cortex for warehouse-native analytics. Multi-model fluency, no single-vendor dependency.

  • Agent-Assisted PR Reviews, Unit Tests & Playwright UI Suites

    Automated PR review agents check style, security vulnerabilities, test coverage, and dependency issues before a human reviewer sees the diff. AI-generated unit tests plus Playwright-driven UI regression suites lifted coverage from under 10% to 40% with no dedicated QA team. Speed and quality, not one or the other.

  • AI-Powered Onboarding & Documentation

    Codebase walkthroughs, architecture Q&A agents, runbook automation, and AI-refreshed documentation that compress new-engineer ramp time 70% - measurable time-to-first-commit on every new hire.

Multi-model AI-augmented developer toolchain
Engineering Practice

Nine Practices That Make AI Work

Tooling is the cheapest and least important part of an AI-native SDLC. DORA's 2025 AI Capabilities Model names seven organizational capabilities that turn adoption into performance: a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control, working in small batches, a user-centric focus, and quality internal platforms. Four of the seven are ordinary engineering discipline. AI amplifies whichever you have. DORA AI Capabilities Model (source dated 2025-11; verified 2026-09)

These are the nine practices I install, in roughly this order. None of them is a product. All of them are how the review queue stays short while generation volume climbs.

1. Context Engineering

Every repo an agent touches has an instruction file at the root: CLAUDE.md for Claude Code, AGENTS.md for Codex, Cursor, Copilot, and most others (Claude Code reads AGENTS.md natively since 09-2026 when no CLAUDE.md exists). Build and test commands, the conventions the code does not make obvious, the risk tiers, what never to touch, and the security rules models get wrong most.

Keep it short. The test for each line is "would removing it cause a mistake?" Instruction files are code - reviewed, versioned, CODEOWNED. They are also advisory. Hooks and permission settings are enforcement; the never-events live there.

2. The Ticket as a Contract

A ticket a machine can test against is the entire unlock. Tests-first, agent implementation, and compliance evidence all depend on it, so the ticket stops being prose: acceptance criteria in Gherkin, a data contract, the rule or requirement ID it implements, the data classification it touches, required audit-log writes, the feature flag, the rollback plan, and the risk tier.

An ambiguity check against the PRD runs before the ticket is accepted, not after the code lands.

3. Plan First, Then Diff

Approve the plan, not the diff. The agent reads the repo and produces an implementation plan - files, interfaces, migrations, flags, test surface, blast radius, rollback, risk tier. A senior human approves or revises it. Only then does implementation start.

Reviewing a plan takes five minutes and catches architectural mistakes. Reviewing a 1,500-line diff takes an hour and catches typos. Plan approval is where senior judgment scales.

4. Tests First, Mutation Scored

Generate the tests from the acceptance criteria before the implementation exists. File them where QA can see them. Then implement until they pass. The suite becomes an independent statement of what the feature was supposed to do, not an artifact of how the model happened to build it.

Put a mutation score gate on the modules that matter: money math, auth, rule engines, adapters. Coverage is not a gate. Mutation score is.

5. Small Batches and PR Size Caps

Average PR size went up 154% in Faros's 2025 telemetry and review time went up with it. Cap PR size in the pipeline: a soft warning at 400 lines, a hard split request at 800 is a reasonable start.

Stacked PRs and a merge queue make small batches cheap. DORA lists working in small batches as a core AI capability for a reason - it is the single practice that most directly counters the review-queue failure.

6. Risk-Tiered Review

When PR volume rises, uniform review collapses. Either the queue backs up or the reviews go shallow, and both look like progress on a velocity chart. Classify every diff against the approved plan plus the CODEOWNERS map, and route it to a tier that matches its blast radius.

Three tiers are enough. The details are in the next section.

7. Merge Queue

A merge queue batches verification, keeps main green, and gives agents a safe place to repair rebase and lockfile conflicts without ever touching a protected branch.

It is the mechanical answer to "more PRs than the pipeline can serially verify," and it is the piece most teams skip until the third time main goes red on a Friday.

8. Model and Tool Pinning

Pin model versions in CI. Log the version on every agent action. A model upgrade is a change and goes through the same review as a dependency bump: run the eval suite, compare, then promote. The same applies to the review agent's rules and the instruction files.

Unpinned means unreproducible, and unreproducible means the defect class you used to catch escapes without anyone noticing the model changed.

9. The Learn Loop

Every escaped defect becomes a rule in the review agent's library. Every repeated task becomes a versioned skill. Every release generates its evidence pack from the ticket contract. Rules, skills, and fixtures are code with tests: a fixture set of sanitized PRs with known issues, run on every rule change and every model change, with precision and recall tracked per rule.

This is the compounding asset. The models are rented and will be swapped. The context is owned, and it gets better at your codebase every quarter.

Review at Volume

Risk-Tiered Review

Review depth should track blast radius, not calendar availability. Every diff is classified against the approved plan and the CODEOWNERS map, then routed to one of three tiers.

Tier 1

Auto-Merge on Green

Scope: docs, copy, tests, flag-off scaffolding, and dependency bumps in the safe class.

Review: auto-merge on green, with a weekly human audit sample so the tier cannot quietly widen.

Agent-authored: yes.

Tier 2

One Human, AI Pre-Review

Scope: standard feature work behind a flag.

Review: one human approver. AI pre-review resolves the mechanical comments - formatting, naming, obvious anti-patterns - before the human opens the PR, so their attention lands on the design.

Agent-authored: yes, labeled.

Tier 3

Two Humans Plus an Architect

Scope: auth, money and eligibility math, PII paths, migrations, rule engines, external adapters to systems of record, and infrastructure.

Review: two humans plus an architect. No AI auto-approve, ever.

Agent-authored: no.

All three tiers feed the merge queue. An AI approval never satisfies a Tier 3 human approval requirement - the moment it does, the one control that separates a review process from a rubber stamp is gone.

Test Theater

Coverage Is Not a Gate. Mutation Score Is.

AI writes tests that look thorough and assert nothing. A generated suite can execute every line, report 90% coverage, go green on the wrong behavior, and prove exactly nothing about whether the feature works. Coverage cannot catch that, because coverage measures what ran, not what was checked.

Mutation testing can. The tool mutates the code - flips a comparison, deletes a branch, changes a constant - and checks whether the tests notice. Stryker covers JavaScript, TypeScript, and .NET; PIT covers Java; mutmut covers Python; cargo-mutants covers Rust. The share of mutants the suite kills is the mutation score, and it is the test-quality metric coverage was always pretending to be.

The pattern that works: generate, mutate, feed the surviving mutants back to the model as the prompt, regenerate. Most mutation tools shipped incremental modes in 2025 and 2026, which makes the loop affordable per PR rather than a quarterly event. The gate goes on the modules that matter, and it doubles as a quality proof for auditors.

  • Coverage and mutation score are not the same number. On Meta's TestGenEval benchmark, GPT-4o reached 35.2% coverage but only an 18.8% mutation score, and the other frontier models were similar. Meta TestGenEval, ICLR 2025 (source dated 2024-10; verified 2026-09)
  • The gap can be extreme. MUTGEN documented LLM-generated suites with "100% coverage but only 4% mutation score." Feeding the surviving mutants back to the model as a prompt fixed most of it, which is why the loop above is a loop. MUTGEN, IEEE TSE (source dated 2025-06; verified 2026-09)
  • Engineers accept generated tests when each one is proven to matter. Meta's mutation-guided hardening produced 571 generated tests and engineers accepted 73% of them, because each test was shown to kill a specific mutant rather than merely raising a coverage number. Meta ACH (source dated 2025-01; verified 2026-09)
  • Mutation score predicts real bug detection, with one catch. A replication across 11 models and 101,000 generated tests found mutation score correlates strongly with real bug detection (r = 0.86) for regression tests, but weakly when the code under test is already wrong. Generate tests from the acceptance criteria, not from the implementation. ISSTA 2026 replication (source dated 2026-07; verified 2026-09)
Measurement

Measure Outcomes, Not Activity

Baseline before adoption. Capture cycle time from ticket start to production deploy at p50 and p85, PR review turnaround and review time relative to PR size, change failure rate, escaped defects per release, coverage plus mutation score on the critical modules, deployment frequency, and a perceived-productivity survey so you can later measure the perception gap. Capture them before the next expansion, because retroactive baselines are storytelling.

A number without a pre-AI baseline is a story, not evidence. If a clean baseline does not exist, say so in every report rather than back-filling one.

Report These

  • Cycle time, ticket to production, p50 and p85 - the outcome
  • Review latency and review time per line - the constraint
  • Change failure rate - held flat is the honest claim
  • Escaped defects by tier - whether the tiering is calibrated
  • Incidents per deploy - the unreviewed-merge signal
  • PRs merged with no human comment, by tier - oversight erosion
  • Rework rate after review - review effectiveness
  • Mutation score on critical modules - test quality
  • Duplication and cross-file reuse trend - deferred maintenance
  • Agent spend per merged PR, per workflow - unit economics

Do Not Report These as Success

  • AI adoption percentage
  • AI-authored lines or PRs as a share of the total
  • Suggestions accepted
  • Tasks completed per developer, on its own

They go up while quality goes down, and they hand leadership the wrong dashboard. Track adoption for coverage and cost planning, never as an outcome.

The full data behind this - Faros, DORA, METR, and the quarterly paradox test - is in Authoring Is Cheap. Verification Is the Constraint.

Results

Velocity Outcomes

Real results from operationalizing AI-native DevEx across growth-stage engineering organizations. No vanity metrics, no AI theater.

5x

Deploy Frequency

AI-native SDLC running Claude Code, OpenAI Codex, GitHub Copilot, and AWS Kiro as first-class GitHub Actions stages. Code-to-release cycle time down 40%.

23%

PR Throughput Gain

AI-assisted code review, test generation, and documentation refresh shipped throughput improvements without sacrificing quality. SonarQube enforces project standards in CI/CD.

70%

Onboarding Time Cut

AI-assisted documentation refresh plus a lead-mentor program. Test coverage lifted from under 10% to 40% with no dedicated QA team.