Product

How we scaled our platform to 50 million agent runs daily

A deep dive into the infrastructure decisions that let us handle massive scale without breaking a sweat — or the bank.

Published on

Written by

Professional portrait of a man with folded arms against a soft gradient background.

David Park

Scenic lake with grassy shoreline, rocky foreground, and mountains under a bright blue sky.

Last month, we hit a milestone: 50 million agent runs processed in a single day. Zero downtime. Average latency under 100ms. And our infrastructure costs actually went down compared to the previous quarter.

This post is a behind-the-scenes look at how we got here.

When we started, like most startups, we ran everything on a single Kubernetes cluster. It worked fine for our first few customers. But as usage grew, we started hitting walls. Cold starts were killing our latency. Scaling was reactive, not predictive. And our AWS bill was becoming a recurring nightmare in our board meetings.

The first big change was moving to a multi-region architecture. We now run inference workloads across 12 regions globally, automatically routing requests to the nearest healthy cluster. This alone cut our p99 latency by 60%. Users in Singapore were no longer waiting for round-trips to us-east-1.

The second change was rethinking how we handle bursty workloads. AI agents are inherently unpredictable. A single customer might go from 100 requests per minute to 10,000 in seconds. Traditional auto-scaling couldn't keep up — by the time new instances spun up, the burst was over.

Our solution was predictive scaling based on historical patterns combined with aggressive warm pooling. We maintain a reserve of pre-warmed instances that can absorb traffic spikes instantly. The system learns each customer's usage patterns and pre-provisions capacity before they need it. It sounds expensive, but it's actually cheaper than reactive scaling because we waste fewer resources on cold starts.

The third change was optimizing our inference layer. We built a custom request batching system that groups similar queries together, maximizing GPU utilization without sacrificing latency. We also implemented speculative execution for multi-step agent workflows — starting likely next steps before the current step completes.

The results speak for themselves. Our infrastructure now handles 50M+ daily runs with 99.99% uptime. Median latency is 42ms. And we're doing it at a cost per request that's 3x lower than when we started.

We'll be open-sourcing parts of this infrastructure in the coming months. If you're interested in early access, drop us a line.

FAQ

Frequently asked questions

Everything you need to know to get started.

Do we have to move our files?

No. Your documents stay where they are, in Drive, Dropbox, Litify, or Filevine. We read from them and keep watching for new ones. Nothing gets migrated into a platform you would have to leave later.

What exactly do we own?

The data layer. The database, the schema, the pipelines, and the full extraction history live in your firm's cloud and belong to you outright. The workflows and agents that run on top are licensed while we work together. You own the record, and you license the tools that read it.

What happens if we stop working with you?

You keep the data layer and everything in it, running in your own environment. The workflows and agents switch off, and you are free to point other tools at the record you own. Nothing has to be exported, because nothing was ever held anywhere else.

Who can see our client data?

Only the people who can see it today. Everything runs in your firm's own cloud, privilege is gated at the database rather than inside an app, every access is logged, and your data is never used to train anyone's model.

Do we have to leave Litify?

No, and we would not recommend it. Litify stays your system of record. We mirror cases, parties, and assignments from it, and we can write extracted facts back into your fields, so the system your team already knows gets better rather than replaced.

How is it priced?

A build fee for the asset you own, then a flat monthly retainer covering operations and the workflows and agents in use. No per-seat fees and no per-question metering. Your bill does not grow because your team used it more.

What happens when the AI gets something wrong?

You catch it, which is the design. Every answer carries the documents it came from, nothing is filed or sent without a person approving it, and every correction improves the system. The attorney remains responsible for the work product, so the platform is built to make supervising it easy.

Does this replace people on our team?

No. It removes the gathering, retyping, and remembering, which was never the job. Judgment, advocacy, and the client relationship stay exactly where they are.

FAQ

Frequently asked questions

Everything you need to know to get started.

Do we have to move our files?

No. Your documents stay where they are, in Drive, Dropbox, Litify, or Filevine. We read from them and keep watching for new ones. Nothing gets migrated into a platform you would have to leave later.

What exactly do we own?

The data layer. The database, the schema, the pipelines, and the full extraction history live in your firm's cloud and belong to you outright. The workflows and agents that run on top are licensed while we work together. You own the record, and you license the tools that read it.

What happens if we stop working with you?

You keep the data layer and everything in it, running in your own environment. The workflows and agents switch off, and you are free to point other tools at the record you own. Nothing has to be exported, because nothing was ever held anywhere else.

Who can see our client data?

Only the people who can see it today. Everything runs in your firm's own cloud, privilege is gated at the database rather than inside an app, every access is logged, and your data is never used to train anyone's model.

Do we have to leave Litify?

No, and we would not recommend it. Litify stays your system of record. We mirror cases, parties, and assignments from it, and we can write extracted facts back into your fields, so the system your team already knows gets better rather than replaced.

How is it priced?

A build fee for the asset you own, then a flat monthly retainer covering operations and the workflows and agents in use. No per-seat fees and no per-question metering. Your bill does not grow because your team used it more.

What happens when the AI gets something wrong?

You catch it, which is the design. Every answer carries the documents it came from, nothing is filed or sent without a person approving it, and every correction improves the system. The attorney remains responsible for the work product, so the platform is built to make supervising it easy.

Does this replace people on our team?

No. It removes the gathering, retyping, and remembering, which was never the job. Judgment, advocacy, and the client relationship stay exactly where they are.

FAQ

Frequently asked questions

Everything you need to know to get started.

Do we have to move our files?

No. Your documents stay where they are, in Drive, Dropbox, Litify, or Filevine. We read from them and keep watching for new ones. Nothing gets migrated into a platform you would have to leave later.

What exactly do we own?

The data layer. The database, the schema, the pipelines, and the full extraction history live in your firm's cloud and belong to you outright. The workflows and agents that run on top are licensed while we work together. You own the record, and you license the tools that read it.

What happens if we stop working with you?

You keep the data layer and everything in it, running in your own environment. The workflows and agents switch off, and you are free to point other tools at the record you own. Nothing has to be exported, because nothing was ever held anywhere else.

Who can see our client data?

Only the people who can see it today. Everything runs in your firm's own cloud, privilege is gated at the database rather than inside an app, every access is logged, and your data is never used to train anyone's model.

Do we have to leave Litify?

No, and we would not recommend it. Litify stays your system of record. We mirror cases, parties, and assignments from it, and we can write extracted facts back into your fields, so the system your team already knows gets better rather than replaced.

How is it priced?

A build fee for the asset you own, then a flat monthly retainer covering operations and the workflows and agents in use. No per-seat fees and no per-question metering. Your bill does not grow because your team used it more.

What happens when the AI gets something wrong?

You catch it, which is the design. Every answer carries the documents it came from, nothing is filed or sent without a person approving it, and every correction improves the system. The attorney remains responsible for the work product, so the platform is built to make supervising it easy.

Does this replace people on our team?

No. It removes the gathering, retyping, and remembering, which was never the job. Judgment, advocacy, and the client relationship stay exactly where they are.

Own your firm's Intelligence.

Start with a four-week assessment. We map your documents and show you what your files can actually tell you. You keep the findings either way.

Own your firm's Intelligence.

Start with a four-week assessment. We map your documents and show you what your files can actually tell you. You keep the findings either way.

Own your firm's Intelligence.

Start with a four-week assessment. We map your documents and show you what your files can actually tell you. You keep the findings either way.