← Back
Case Study

How I Built a $20M AI Business
Inside a Fortune 500

$2M → $20M Annualized Revenue
500+ Cloud Services Launched
3,000 Monthly Active Users (GenAI)
+22% Pricing Engine Margin Lift

The Problem

Utilities don't think like technology companies. National Grid is a $20B+ business that moves electrons across the northeastern United States and United Kingdom — its core product is infrastructure, its operating model is regulatory, and its organizational culture was built around physical assets and long planning cycles.

When I joined in 2022, "the cloud" was still being sold internally as a cost story. AI was a PowerPoint concept. Digital services — the idea that National Grid could sell technology products to other utilities and municipalities — existed as a $2M line item that nobody quite knew what to do with.

The opportunity was real. Utilities across the country face the same infrastructure modernization challenges National Grid had already solved: grid management software, outage analytics, customer experience tooling, compliance automation. Every solution we'd built for ourselves was a potential product. We just needed a business model, an operating model, and a PM who'd actually ship things.

What I Built

Over three years, I led the fixed-price digital services business from $2M to $20M annualized revenue. The key word is "fixed-price" — we weren't doing time-and-materials consulting. We were packaging our existing infrastructure expertise into defined, scoped products with predictable delivery timelines and outcomes. That's a fundamentally different motion than what utilities typically buy.

The cloud portfolio grew from zero to 500+ services. Some were internal: infrastructure modernization, developer tooling, data platforms. Many became external products we sold to other utilities. The distinction matters because it forced a discipline that most internal platform teams avoid — you had to define what "done" actually meant.

In 2024, we launched an internal GenAI platform that reached 3,000 monthly active users. This wasn't a pilot with 50 power users. It was production, with real data, real workflows, and real accountability to business outcomes. The hardest part wasn't the technology. It was building the operating model that let people actually use it safely and at scale.

I scaled the cloud and data organization from 5 to 35 engineers, promoted from Technical PM to Senior PM in 18 months, then to Principal PM in 6 months. The promotions mattered less than what drove them: shipping things that worked.

The Operating Model

Most companies deploy AI wrong. Not technically wrong — the models work fine. Operationally wrong. They treat AI as a feature when it needs an operating model.

Five failure modes I saw repeatedly across the utilities we worked with:

  1. The procurement gate: The approval process for AI tools was designed for software licensing, not for systems that generate output. Weeks of legal review for a chatbot pilot, while the actual data governance question — what can this system see, and what can it say? — went unasked.
  2. Model-as-a-feature thinking: "We're adding AI to our billing system." What that actually means is undefined. Which inputs? What output format? Who validates? What happens when it's wrong? Companies shipped the model before they shipped the process.
  3. No eval framework: The demo worked. The pilot worked (with 10 friendly users). The production rollout surfaced edge cases nobody had tested because there was no systematic way to test them. Evals aren't optional; they're the product.
  4. No cost model: Token costs are a real operating expense, and they scale with usage in ways that differ from traditional software. Teams that launched without a cost model discovered this at month-end. The discipline of "how much does it cost to serve one user for one month" was completely absent.
  5. No governance layer: Who decides when the model is wrong? Who can turn it off? Who reviews its outputs systematically? Without this, you're running a system with no ownership and no accountability.

What worked at National Grid was treating AI deployment as a product discipline, not a technology project. Every GenAI capability we shipped had: a defined scope of what it could and couldn't do, an eval framework before launch, a cost model, a governance owner, and a rollback path. Boring. But it's why we had 3,000 active users instead of a pilot that died after the executive sponsor changed jobs.

The AI Pricing Engine

The most technically interesting thing I built was the dynamic pricing engine for fixed-price digital services. The core problem: pricing utility technology products is hard because the inputs are heterogeneous (different grid sizes, different data maturity, different existing infrastructure), the market comparables are sparse, and historical deal data was locked in sales spreadsheets.

The engine combined several input layers: customer infrastructure data (grid complexity, service territory size, existing tooling), historical deal outcomes (what we'd charged, what we'd delivered, what the margin had been), and real-time market signals. An LLM layer was used for structured extraction from unstructured deal notes and for reasoning about customer-specific context — not for the pricing decision itself.

The lesson: LLMs are excellent at turning messy human-generated data into structured inputs for a deterministic model. They're poor substitutes for the deterministic model. Use them in the right layer.

The pricing model itself was a gradient boosting ensemble with interpretable feature importance — meaning a sales rep could understand why a deal was priced a certain way and push back with domain knowledge. That feedback loop was the actual moat, not the model. The model improved because salespeople corrected it.

Result: 22% margin improvement on deals that used the engine versus deals that didn't, measured over a 12-month production window. The honest caveat: selection effects exist. Reps used the engine more on deals they were confident about. We controlled for this partially but not perfectly.

What I'd Do Differently

Three things:

Start the eval framework on day one, not after the pilot. We built our GenAI eval framework reactively — after we'd already launched and discovered edge cases in production. The eval work was unglamorous and took longer than it should have because we were retrofitting it onto a system that wasn't designed for systematic testing. Evals should be a launch criterion, not a post-launch remediation.

Hire a data scientist earlier. I was doing too much of the analytical work myself, which meant product decisions moved slower than they should have. A dedicated person who owned the model quality and the eval process would have been multiplicative, not additive.

More aggressive on the governance model. We got the governance right by the end, but the early months were more ad-hoc than they should have been. In hindsight, having a formal AI governance structure from launch — with explicit policies about what the system could do, who owned it, and how exceptions were handled — would have accelerated adoption and reduced the compliance conversations we had to have retroactively.

The Playbook

Five things that actually matter when deploying AI at enterprise scale:

  1. Define the operating model before the technology. Who owns it? Who can turn it off? What can it see? What can it say? What happens when it's wrong? These are not technology questions. Answer them first.
  2. Evals are the product. A model without a systematic way to measure its quality is not a product — it's a prototype. Your eval framework is what separates a pilot from something you can stand behind at scale.
  3. Ship the cost model with the system. Token costs, inference infrastructure, human review time — know what it costs to serve one user for one month before you launch. This changes product decisions in ways that improve the business.
  4. Use LLMs in the right layer. Structured extraction, reasoning about unstructured context, interface generation — LLMs excel here. Deterministic business logic, numerical optimization, compliance rules — keep those deterministic. The architecture is about knowing which layer you're in.
  5. The feedback loop is the moat. Whatever your system does, design the feedback loop first. The model that improves because real users correct it will outperform the model that doesn't get corrected, regardless of where they start.
Connect on LinkedIn ← Back to Portfolio