Pitstop

Five AI agents. One pit crew.

A digital pit stop for your website.

Bring your site in. A crew of AI agents measures it, finds what is slowing it down, writes the fix, and proves it worked — then hands you a pull request. You wave it through.

Install the App Sign in

Nothing to configure. Nothing to learn. Install it on one repository and the next push gets a crew. Three to eight minutes, start to pull request.

run 4nkpv lazydevpro/redline-demo-site

A run on the route /about: LCP fell from 10.1 seconds to 3.9 seconds, 61% faster, and the change was opened as pull request 10.

LCP 3.9s CLS 0.012 INP 0ms SAMPLES 5 VERDICT shipped

Judged on the clock

Your page is the car.

A lap time is not an opinion. Every fix is run twice, and if the second run disagrees with the first, nothing ships.

Five loads per measurement Median reported, never the best lap Re-measured after every patch

Meet the crew

Five agents. One job each.

The
Surveyor

Puts your site on the clock

The
Attributor

Finds the file behind the number

The
Surgeon

Writes the fix, one at a time

The
Referee

Cannot see how the fix was made

The
Scribe

Writes the pull request

The problem

Everyone measures.
Almost nobody fixes.

Your site has been telling you it is slow for months. A score is not a fix — it is a chore, assigned to nobody, and it loses every sprint.

You do not need another dashboard. You need the diff.

01

Red numbers nobody opens

A metric that goes red and stays red stops being information. By week three it is a muted Slack channel.

02

Advice you cannot merge

"Serve images in next-gen formats" never says which of your 340 images, or what fixing them would be worth.

03

Fixes nobody re-checks

The change ships and is never measured again. Plenty do nothing at all. Some quietly make things worse.

How a stop works

In. Worked on. Out.

Every push is a stop. The end of it is a pull request, because shipping is still your call.

On the clock

Built, served and timed on every route. Five loads each, and we report the median — never the flattering one.

Find the fault

The slow number is traced to the file behind it. No confident answer, no patch.

Make the change

A real diff — re-encoded images, <picture> wrappers, dimensions, fetchpriority, defer, font-display.

Back on the clock

Measured again from scratch. This is the step that separates a fix from a hope, and the one everyone else skips.

Ship it, or bin it

Six checks, all of which must pass. Fail one and it is binned, with the reason kept. Pass them all and the pull request lands.

The six criteria

  • Target metric improved ≥ 10%
  • No other Core Web Vital worse by > 5%
  • Layout did not move
  • Build stayed green
  • No new console errors or failed requests
  • Samples agreed closely enough to compare

What you get

Faster pages you never had to chase.

All of it shipping today. None of it a roadmap item.

Six kinds of fix, one pull request

Bloated images, CSS backgrounds, unsized media, blocking scripts, fonts that hide your text. Four problems, one diff, one review.

The failures are kept too

Every rejected patch is saved with the check it failed. Audit it instead of trusting it.

See it, do not just read it

Both builds captured side by side. Proof your page still looks right.

It would rather do nothing

Every fix is guarded — anything already handled is left alone. No patch aimed at the wrong file.

Accessibility flagged, never guessed

Findings arrive with their WCAG criterion, and no agent writes alt text. Nothing plausible and wrong in your codebase.

Watch the stop happen

A live timeline while the crew works, failed attempts included. No log reading required.

Measured, not claimed

Three real runs. One of them refused.

From lazydevpro/redline-demo-site, on Cloud Run. Every run and every rejected attempt is in that repository.

10.1s → 3.9s /about — LCP, 61% faster. Hero image re-encoded and downscaled, image-set() background with the original kept as fallback, font-display: swap. Shipped as pull request #10.
3.76 MB → 1.69 MB /gallery — five photographs re-encoded and resized. LCP 22.7s to 11.1s on the same run.
Withdrawn /pricing — the patch made LCP 4% better. The bar is 10%, so the attempt was thrown away and nothing was opened. This one is on the page on purpose.

FAQ

The questions you should be asking.

Still unanswered? The security and data page goes further, including the parts that are not finished yet.

AI writing code straight into my repository. Why would I allow that?

Because it does not get the final say — you do, and neither does the agent that wrote the change.

The Surgeon writes a patch. The Referee then rebuilds your site, re-measures it five times, compares every screenshot, and decides. The Referee cannot see the Surgeon's work — the reasoning is not in its context, so there is no argument to be persuaded by, only numbers. Most patches lose that argument and are thrown away.

What reaches you is a pull request that already survived a build, a re-measurement and a visual diff. It still needs your approval, and nothing merges without it.

Which models are you using?

Gemini, through Google Vertex AI. Five agents share the model but not the job: Surveyor, Attributor, Surgeon, Referee and Scribe, each with its own tools and its own authority to stop the run.

The measuring is not AI at all — that is Lighthouse and Playwright on fixed hardware. Agents decide what to try and whether it worked; the stopwatch is a stopwatch.

What does it actually change in my code?

Markup and CSS, plus image binaries. Concretely: a <picture> wrapper around an <img>, width and height attributes, fetchpriority="high", removal of loading="lazy" from an above-the-fold hero, defer on a blocking script, font-display: swap, an image-set() background declaration, and re-encoded WebP files alongside the originals.

It does not touch your application logic, your framework config, your dependencies, or your build pipeline.

Will it open a pull request on every push?

No. It opens one only when a patch passed all six criteria, and it reuses the same branch per route — so a second run that improves the same page updates the existing pull request instead of opening another.

Most runs on a healthy site open nothing, which is the correct outcome and shows in the dashboard as a run with no findings.

How do you know the change actually helped?

Because the site is measured twice — once before the patch and once after — with the same routes, the same five samples, and the same machine class. The pull request body carries both numbers.

Improvement is judged on the median of five loads, not a single run, and a metric has to move by at least 10% to count. Anything smaller is inside the noise this kind of measurement produces.

What if the fix makes something else worse?

The attempt is withdrawn. No other Core Web Vital may degrade by more than 5%, the rendered layout is compared box by box, the build must still succeed, and the page must not log any console error or failed request it did not log before.

That last criterion exists because a patch once "improved" every metric by breaking a click handler — less JavaScript ran, so the page got faster and stopped working.

What access does it need, and what leaves my machine?

Contents read & write (to clone, and to push a redline/* branch), Pull requests write, Checks write, and Metadata read. Nothing else.

Your repository is cloned into a container that exists for one run and is destroyed after it. Your build command runs there, because measuring a site means building it. Screenshots and Lighthouse reports live in a private bucket for 90 days so the dashboard can show them. Nothing is sent to any third party — no analytics, no telemetry.

Which frameworks does it support?

Anything that builds to static files. Detection needs a build script in package.json — it then runs npm install && npm run build — so Next.js static exports, Astro, Vite, SvelteKit static and Eleventy work without configuration. The output directory is read from your Vite config if you have one, and otherwise guessed from dist, build, out, public, _site.

If none of that fits — a site with no package.json, a different package manager, or an output directory somewhere else — add a redline.config.json with a buildCommand and outputDir and it will use those instead of guessing. Without either, the run stops and tells you so rather than picking something.

Server-rendered apps that cannot be built and served statically are not supported yet.

How long does a run take, and what does it cost me?

Three to eight minutes depending on how many routes it measures, starting within seconds of your push. Two runs never race on one repository — the second is skipped rather than queued behind a ten-minute measurement.

Can I look at it without installing anything?

Yes. The demo repository is public and holds every pull request Pitstop has opened against it, including the bodies with before-and-after tables, and the branches for attempts that were withdrawn.

Push something slow.
See what comes back.

One repository is enough. Your next push gets a crew, and you get either a pull request with the numbers in it — or a straight answer about why there wasn't one.

Install the App Read a real pull request

ROUTE /about LCP 3.9s CLS 0.012 INP 0ms SAMPLES 5 BUDGET 2.5s VERDICT SHIPPED