Type check green. Sixty-plus unit tests green. Local end-to-end journey green. Then you point the thing at a real domain with real keys and four separate things fall over, none of which a local test suite was ever going to catch.
The app is PrintLatch, a quote-and-approval tool for screen printing shops. It’s ours, so the failures below are ours too. The stack is Cloudflare Workers, D1, R2 and Stripe subscriptions — but three of the four findings have nothing to do with that stack specifically. They’re shapes of bug you’ll hit on any launch.
1. Tax ID collection rejects every checkout for an existing customer
Symptom: every single Checkout session creation returned a 400. The API error:
Tax ID collection requires updating business name on the customer.
To enable tax ID collection for an existing customer,
please set `customer_update[name]` to `auto`.
The fix is one field:
customer_update: { name: "auto", address: "auto" },
automatic_tax: { enabled: false },
tax_id_collection: { enabled: true },
Why it’s easy to miss: this only fires when you pass an existing customer to
Checkout. If you let Checkout create the customer itself, you never see it. Plenty of
apps create the Stripe customer at signup — as soon as you do that and turn on tax ID
collection, every checkout dies.
The deeper lesson is about where it surfaced. Our test suite mocks the Stripe client, so it validated our logic perfectly and told us nothing about what Stripe would accept. A mocked provider tests your code. It does not test your contract with the provider. The only thing that finds this is a real API call to a real account.
2. A retry that can never succeed, retried forever
This one is subtler and it’s the one worth taking away.
Creating a Checkout session is a network call that can fail after Stripe has already done the work. So the standard defensive pattern is: persist the exact parameters and an idempotency key before the call, and if you find a leftover pending attempt later, replay it with the same key. Stripe returns the original session instead of making a second one. Good pattern. We had it.
The hole: it assumes every failure is a lost response. It isn’t. When Stripe rejects the request — a 400 from bad parameters, or a card error — replaying identical parameters under the identical key can never produce a different answer. And because the recovery step runs first, before anything else, the studio is now permanently wedged. Every future checkout attempt starts by trying to recover the dead one, dies on the same 400, and returns a generic error. There is no path out through the UI.
That’s how finding #1 turned from a bug into a trap: the tax ID error left a poisoned attempt behind, and even after we fixed the parameters, the stored attempt kept being replayed with the old ones.
The fix is to classify the failure rather than treat them all alike:
function rejectedOutright(e) {
// An idempotency conflict means the original request is still in flight — recoverable.
if (e?.rawType === "idempotency_error") return false;
return (
e?.type === "StripeInvalidRequestError" ||
e?.type === "StripeCardError" ||
e?.rawType === "invalid_request_error" ||
e?.rawType === "card_error"
);
}
Rejected outright means retire the stored attempt and let the next call build fresh parameters. Everything else — timeouts, 5xx, rate limits, idempotency conflicts — stays recoverable exactly as before. Two regression tests now cover it: one asserts a rejected attempt gets retired, and one asserts a stale poisoned attempt doesn’t block a new checkout on a different plan.
If you have retry logic anywhere, ask it one question: is there a failure here that retrying can never fix? If yes, and you retry it first thing on every subsequent request, you have built a permanent wedge.
3. A safe default that quietly argues with your sitemap
Our root layout sets robots: { index: false, follow: false }. That’s the correct
default — the dashboard, billing, settings and customer-approval pages must never be
indexed, and a safe default means a new private page can’t leak by forgetting a line.
Public marketing pages override it. Except three didn’t: contact, terms and privacy had
no metadata of their own, so they inherited noindex — while being listed in the sitemap.
A sitemap is an assertion that a URL is worth indexing. A noindex on that URL is the
opposite assertion. Search Console flags the conflict, and the pages don’t get indexed.
What caught it wasn’t reading the config. It was crawling our own sitemap and checking each URL’s actual response:
curl -s https://example.com/sitemap.xml | grep -o '<loc>[^<]*' | sed 's/<loc>//' |
while read u; do
echo "$u $(curl -s "$u" | grep -c 'content="noindex')"
done
Nineteen URLs, three of them lying. Keep the safe default; just assert the exceptions explicitly, and verify against the deployed site rather than the source.
While you’re there, one more thing worth knowing: IndexNow does not reach Google. Bing, Yandex, Seznam and Naver participate; Google does not. If you submit to IndexNow and watch for Google traffic, you’ll conclude something is broken when nothing is. Google is Search Console, full stop.
4. A codemod that wasn’t idempotent
We have a one-time script that walks JSX and wraps untranslated text in a <T> component.
It skips files that already import T. The guard:
if (source.includes('import { T }')) continue;
That matches import { T } from "@/app/i18n". It does not match
import { T, LanguageSwitch } from "@/app/i18n", and it does not match
import { T } from "./i18n". So re-running the script on a finished codebase produced
<T><T>Pricing</T></T> and duplicate imports across twenty-three files.
Two rules fall out of this:
- A skip-guard must match every shape of the thing it’s looking for. Substring matching on source code is almost always too narrow. A regex against the import specifier is the right granularity here.
- A codemod’s real test is running it twice. If the second run isn’t a no-op, it isn’t
safe to keep in the repo — and someone will run it again, because it’s sitting right
there in
scripts/.
The related trap: the script also rewrote its output manifest on every run. Once the guard was fixed and every file was correctly skipped, it wrote an empty manifest, because “nothing new was extracted” and “nothing exists” looked identical. Incremental tools should merge into their output, not overwrite it.
Bottom line
None of these four were caught by types, unit tests, or a local end-to-end run. What caught them was pointing the real thing at the real world: one live API call against a real Stripe account, and one crawl of our own deployed sitemap.
Budget time for that step. It is not a formality after the build — it’s a different class of test, and it’s the only one that checks the assumptions you didn’t know you were making.