Or press ESC to close.

Delivered Twice, Charged Twice: Testing Webhooks with Playwright

Oct 11th 2026 19 min read
medium
javascriptES6
playwright1.64.0
nodejs24.13.0
api
security

With most API tests, you decide when the request is sent, what it contains, and how many times it happens. Webhooks take all three decisions away from you. The sender picks the timing, retries on its own schedule, and makes no promise about the order events arrive in, so the same event can reach your server twice and a later event can show up before an earlier one. A receiver that passes a happy-path test can still charge a customer twice in production. This post builds a small billing webhook receiver and tests it with Playwright's request fixture, covering signature verification, replayed requests, duplicate deliveries, out-of-order events, and retries after a failure.

Why Webhooks Break Differently

In a normal API test, your code is the client. You build the request, send it, and assert on the response, and every part of that exchange is under your control. A webhook flips the roles. Your server becomes the API, and the client is someone else's system: a payment provider, a shipping service, a source control host. It has its own delivery rules, and those rules are written to protect the sender, not your database.

The first rule is that delivery is at-least-once. If your endpoint does not answer with a success status quickly enough, the sender assumes the event was lost and sends it again. That includes the case where your server received the event, did all the work, and was simply slow to reply. From the sender's side that looks identical to a failure, so a duplicate arrives for an event you already handled. Duplicates are not an edge case here. They are how the system is designed to behave, and any side effect your handler triggers, like sending an email or creating an invoice, will run twice unless you prevent it.

The second rule is that nobody promises ordering. Events are delivered in parallel, and failed ones are retried later, so a subscription.updated event can land before the subscription.created event it depends on. Imagine the first delivery of the older event fails, the newer one goes through, and then the retry of the older event arrives. If your handler simply writes whatever it receives, it will quietly roll the record back to a state that is already out of date.

The third problem is that the endpoint has to be public. The sender needs to reach it from the internet, which means anyone else can reach it too. The only thing separating a real event from a forged one is a signature computed with a secret you share with the sender. Even a genuine, correctly signed request can be captured and sent again later, so a valid signature alone does not prove the request is fresh.

What makes all of this easy to miss is how webhook handlers usually get tested. A test posts one valid payload, sees a 200, and moves on. That test passes for a receiver that ignores signatures, applies every duplicate, and trusts whatever order it is given. None of those failures appear until the conditions that cause them do, and by then they show up as a customer charged twice, two welcome emails in one inbox, or an account stuck on the wrong plan.

The good news is that every one of these behaviors can be tested from the outside. The sender is just an HTTP client that sets a few headers and a body, and a test can play that role. Playwright's request fixture is built for exactly this, since it can post a signed payload, post it again, post it with a stale timestamp, or post five copies at the same time. To do that, we first need something to test.

The Example: A Small Billing Receiver

The receiver used throughout this post is a small Node.js server with no runtime dependencies. It exposes one endpoint, POST /webhooks/billing, and pretends to be the billing side of a subscription product. When a subscription is created, it stores the subscription and sends a welcome email. When one is updated, it changes the stored status. The email is the important part. It is a side effect you can count, and counting it is how the tests will later prove that a duplicate delivery did nothing.

Every event arrives as a JSON body with an id, a type, and the subscription data. The version field is a number that goes up each time the subscription changes, and it is what lets the receiver tell a newer event from an older one:

                
{
  "id": "evt_1",
  "type": "subscription.created",
  "data": { "subscriptionId": "sub_1", "status": "active", "version": 1 }
}
                

The sender also attaches two headers: x-webhook-timestamp, which is the Unix time in seconds, and x-webhook-signature. The signature is an HMAC-SHA256 hash computed with a secret that only the sender and the receiver know. What gets hashed is the timestamp, a dot, and the raw request body, so changing either one produces a different signature. The sender and the receiver both use this one small function, and so will our tests when they play the sender:

                
export function sign(secret, timestamp, rawBody) {
  const digest = createHmac('sha256', secret)
    .update(`${timestamp}.${rawBody}`)
    .digest('hex');
  return `sha256=${digest}`;
}
                

Putting the timestamp inside the signed string is what makes replay protection possible later. If the timestamp sat outside it, an attacker could take a captured request and attach a fresh time without invalidating the signature. With it inside, a request is only valid for the moment it was signed.

When a request arrives, the receiver runs four checks in a fixed order, and each one has its own way of failing:

Step Check If it fails
1 Signature matches the raw body 401, nothing is processed
2 Timestamp is within five minutes 401, nothing is processed
3 Event ID has not been seen before 200 with a "duplicate" status, nothing is processed
4 Processing succeeds 500, so the sender retries

A few of those responses are deliberate choices. A duplicate gets a 200 rather than an error, because from the sender's point of view the event was delivered and it should stop retrying. A processing failure gets a 500 for the opposite reason, since a success status would tell the sender to give up on an event we never applied. The order matters too. The signature is verified against the raw body before the JSON is even parsed, so an unauthenticated request never reaches any of the logic below it.

The last piece is the part that does the actual work. It lives in a small in-memory store and contains the two rules the ordering and duplicate tests will lean on. State only moves forward, meaning an event is applied only if its version is higher than the one already stored. The welcome email is recorded whenever a creation event is processed, which gives us a list to inspect:

                
if (!current || version > current.version) {
  subscriptions.set(subscriptionId, { status, version });
}

if (event.type === 'subscription.created') {
  emailsSent.push(subscriptionId);
}
                

Two things in this example are simplified on purpose. The store keeps everything in memory, where a production receiver would use a database, and the receiver remembers seen event IDs in a Set rather than a table. Both are fine for a demonstration and both are worth keeping in mind, and we will come back to them at the end of the post. For now, the receiver is small enough to read in one sitting, and every behavior in the table above is something we can test.

Setting Up Playwright's request Fixture

Playwright is usually associated with browsers, but its test runner works just as well without one. The request fixture gives every test an HTTP client, and a spec file that only uses it never launches a browser. That means @playwright/test is the only dependency the suite needs, there are no browser binaries to download, and the tests start in milliseconds. The config is correspondingly small, with fullyParallel turned on so tests run side by side.

Before any request can be sent, the receiver has to be running. Instead of starting it once for the whole suite, the tests get their own through a custom fixture. Fixtures in Playwright can do setup, hand something to the test, and then clean up after it, all in one function. This one creates a fresh store, starts the receiver on port 0, which tells the operating system to pick any free port, and hands the test the URL and the store. When the test ends, the server is closed:

                
receiver: async ({}, use) => {
  const store = createStore({ delayMs: 20 });
  const server = createReceiver({ secret: SECRET, store });
  await new Promise((resolve) => server.listen(0, resolve));

  await use({
    url: `http://localhost:${server.address().port}/webhooks/billing`,
    store,
  });

  await new Promise((resolve) => server.close(resolve));
},
                

There are two reasons this is set up per test rather than once per suite. The first is isolation. Every test gets an empty store and an empty list of seen event IDs, so a test that processes evt_1 cannot turn the next test's delivery of the same event into a duplicate. Because each server also has its own port, the tests are free to run in parallel. The second reason is access. The receiver runs in the same process as the test, so the test can read the store directly and assert on the real outcome, such as how many welcome emails were sent, instead of inferring it from an HTTP response. The delayMs of 20 milliseconds adds a short pause inside the processing step. It does not change what the receiver does, but it widens the window in which two simultaneous deliveries overlap, which is what the concurrency test will need.

The second fixture is the one that plays the sender. A real provider takes an event, signs it, and posts it, and our tests need to do the same thing while also being able to get it wrong on purpose. The function takes the event plus an optional set of overrides, and every override defaults to the correct value. A test that passes nothing sends a perfectly valid request, and a test that wants to break exactly one thing overrides exactly one option:

                
const {
  secret = SECRET,
  timestamp = Math.floor(Date.now() / 1000),
  body = JSON.stringify(event),
  signedBody = body,
} = options;
                

The last option is the subtle one. body is what gets sent over the wire, and signedBody is what the signature is computed over. They are the same by default, but a tampering test can set them apart to simulate a request that was altered after it was signed. Sending the request is then a single call to the request fixture:

                
return request.post(receiver.url, {
  headers: {
    'content-type': 'application/json',
    'x-webhook-timestamp': String(timestamp),
    'x-webhook-signature': sign(secret, timestamp, signedBody),
  },
  data: body,
});
                

One detail here is easy to get wrong. The data option is given a string, not an object. When Playwright receives an object it serializes it itself, which leaves the exact bytes on the wire out of the test's hands. A string is sent exactly as provided, so the body the receiver hashes is the body the test signed, and a test can deliberately send something different. A signature is computed over bytes, so the bytes are what we want to control. The function also returns the response untouched and asserts nothing. Each test decides what a correct outcome looks like, which keeps the helper reusable for the success cases and the failure cases alike.

With both fixtures in place, a test reads almost like a sentence. It names the event it wants to send, optionally breaks one thing about the request, and then asserts on the status code and on what the store ended up holding. The first thing to try breaking is the signature.

Test 1: Reject Anything You Can't Verify

The signature check is the front door of the receiver, so it gets tested first. The goal is to prove two things: a genuine request gets in, and a request that cannot be verified does not. The order matters, because the negative tests only mean something once the positive one passes. A receiver that rejects every request with a 401 would pass all the "reject" tests while being completely broken.

So the first test is the control. It builds a creation event with the buildEvent helper, sends it with no overrides, and checks two things. The status is 200, and the subscription now exists in the store with the status and version the event carried:

                
const created = buildEvent({ id: 'evt_1', type: 'subscription.created', status: 'active', version: 1 });

test('accepts a correctly signed event', async ({ deliver, receiver }) => {
  const res = await deliver(created);

  expect(res.status()).toBe(200);
  expect(receiver.store.subscriptions.get('sub_1')).toEqual({ status: 'active', version: 1 });
});
                

The second test covers the most direct attack on a webhook: changing the payload in transit. The request is signed over the original body, then sent with a modified one in which the subscription status has been switched to canceled. This is where the body and signedBody options from the previous section come in. The signature on the wire is perfectly valid for the original event and wrong for the one that was actually sent:

                
const tampered = JSON.stringify({ ...created, data: { ...created.data, status: 'canceled' } });

const res = await deliver(created, { body: tampered, signedBody: JSON.stringify(created) });

expect(res.status()).toBe(401);
expect(receiver.store.subscriptions.size).toBe(0);
                

Notice that the test asserts on the store as well as the status code. A 401 on its own only tells you what the receiver said. The empty store tells you what it did, and those are not always the same thing. A handler that responded with an error but still applied the event would pass a status-only check and quietly corrupt data.

The third test covers a different failure. Here the body is untouched, but the signature was produced with a secret the receiver does not know, which is what a request from someone who found the URL and guessed at the format would look like. Everything about the request is well-formed except the one thing an outsider cannot fake:

                
const res = await deliver(created, { secret: 'whsec_attacker' });

expect(res.status()).toBe(401);
expect(receiver.store.subscriptions.size).toBe(0);
                

Together these three tests pin down the behavior from both sides. A receiver that parses the JSON and trusts it would pass the first test and fail the other two. One that rejects everything would fail the first. Only a receiver that actually checks the signature against the exact bytes it received passes all three. What these tests cannot see is how the comparison is done. The receiver uses a constant-time comparison so response timing does not leak information about the signature, but that is a property of the implementation and not something a status code reveals, so it belongs in code review rather than in a test.

There is still a gap, though. Every request so far was sent at the moment it was signed. A request with a perfectly valid signature can also be dangerous, if it is old.

Test 2: Block Replayed Requests

A signature proves that a request was created by someone who holds the secret. It says nothing about when. If an attacker captures one genuine delivery, from a log, a proxy, or a compromised machine along the way, they can send those exact bytes again tomorrow, next week, or next year, and the signature will still verify. This is a replay attack, and the signature check alone has no defense against it.

The defense is the timestamp, and this is why it was placed inside the signed string earlier. The receiver compares the timestamp in the header to the current time and refuses anything outside a five-minute window. Because the timestamp is part of what was signed, an attacker cannot refresh it without invalidating the signature. The check is only three lines:

                
const ageSeconds = Math.abs(now() / 1000 - Number(timestamp));
if (!(ageSeconds <= toleranceSeconds)) {
  return reply(res, 401, { error: 'timestamp outside tolerance' });
}
                

Two small details are worth noticing. Math.abs means a timestamp too far in the future is rejected just like one too far in the past. And the condition is written as !(age <= tolerance) instead of age > tolerance. If the header holds something that is not a number, the age becomes NaN, every comparison with it is false, and the first form rejects the request while the second would let it through.

Testing this looks different from the previous tests, because the request must be valid in every way except its age. The deliver fixture helps here, since it signs with whatever timestamp it is given. Passing an old one produces a request with a correct signature for a moment ten minutes in the past, which is exactly what a replayed capture would look like:

                
const tenMinutesAgo = Math.floor(Date.now() / 1000) - 600;

const res = await deliver(created, { timestamp: tenMinutesAgo });

expect(res.status()).toBe(401);
expect(receiver.store.subscriptions.size).toBe(0);
                

This is what makes the test trustworthy. Because the signature is valid, the only thing that can cause the 401 is the freshness check. If the test had used a bad signature as well, it would pass for the wrong reason and tell you nothing about replay protection. Isolating one variable per test is the same discipline as in the previous section, and it is what lets a failure point directly at the check that broke.

It is just as important to understand what this test does not cover. A timestamp window shrinks the replay opportunity from forever to five minutes, but it does not close it. An attacker who replays a captured request within that window passes both the signature check and the freshness check. Providers also deliver real duplicates inside that window all the time, for reasons that have nothing to do with attackers. Telling a repeated event apart from a new one is a different problem, and it needs a different kind of check.

Test 3: Handle Duplicate Deliveries

This is the test the article is named after. When the same event reaches the receiver twice, the second delivery must be acknowledged and then ignored. Handling it is called idempotency: doing the operation again leaves the world in the same state as doing it once. In our example the visible proof is the welcome email. A customer should receive exactly one, no matter how many times the sender delivers the event.

The mechanism is simple. Every event carries a unique id, and the receiver keeps a set of the ones it has already taken responsibility for. If an incoming ID is in the set, it replies with a success status and does nothing else. If not, it adds the ID and carries on to process the event:

                
if (claimed.has(event.id)) {
  return reply(res, 200, { status: 'duplicate' });
}
claimed.add(event.id);
                

The first test is the straightforward case. It sends the same event twice, one after the other, and checks three things: the first response says the event was processed, the second says it was a duplicate while still returning a 200, and only one welcome email exists. The status code on its own is not enough here. The email list is what proves the second delivery had no effect:

                
const first = await deliver(created);
const second = await deliver(created);

expect(await first.json()).toEqual({ status: 'processed' });
expect(second.status()).toBe(200);
expect(await second.json()).toEqual({ status: 'duplicate' });
expect(receiver.store.emailsSent).toHaveLength(1);
                

If this were the only duplicate test, the receiver would look finished. But sequential duplicates are the easy case. In production the more dangerous duplicates arrive at the same moment. A sender that times out and retries while the first request is still being processed produces two requests that are both in flight at once, and a check-then-act sequence has a gap in it. Both requests ask whether the ID has been seen, both are told no, and both go on to send the email.

The second test is built to find that gap. It fires five identical deliveries at the same time with Promise.all, and the 20 millisecond delay in the store from the fixture keeps all five overlapping. The assertions are strict: every request gets a 200, exactly one of them is reported as processed, and exactly one email exists:

                
const responses = await Promise.all(Array.from({ length: 5 }, () => deliver(created)));
const bodies = await Promise.all(responses.map((r) => r.json()));

expect(responses.every((r) => r.status() === 200)).toBe(true);
expect(bodies.filter((b) => b.status === 'processed')).toHaveLength(1);
expect(receiver.store.emailsSent).toHaveLength(1);
                

The reason the receiver passes comes down to where the claim happens. The has and add calls in the first snippet sit next to each other with no await between them, and they run before the slow processing step. Node handles one piece of synchronous code at a time, so a second request cannot slip in between the check and the claim. If the ID were only recorded after processing finished, all five requests would pass the check before any of them had recorded anything.

That is not just a theory. When I moved the claim to after the processing step to test exactly this, the sequential test kept passing and only the simultaneous one failed. A suite with just the first test would have called that broken receiver correct. This is the real value of the concurrent test, and it is why a duplicate test that never sends two requests at once is incomplete.

One caveat applies. A set in memory protects a single process for as long as it runs. A real receiver behind several instances, or one that restarts, needs the same guarantee from shared storage, usually a unique constraint on the event ID in the database. We will return to that at the end. First, there is a failure that duplicates are often confused with: events that arrive in the wrong order.

Test 4: Survive Out-of-Order Events

Idempotency answers the question "have I seen this event before?" It has nothing to say about a different problem, where every event is new but they show up in the wrong sequence. Each of these events has its own ID, so the duplicate check lets all of them through, and a handler that applies whatever it receives ends up with the oldest event written last.

Here is how that happens in practice. A subscription is created and then updated a moment later. The delivery of the creation event fails on the first attempt and is scheduled for a retry. The update goes through immediately. When the retry finally lands, the receiver sees a creation event for a subscription that has already moved on, and a handler that simply writes it would put the record back to its original state. The customer's subscription is now wrong, and nothing in the logs looks like an error.

Arrival time cannot be trusted to say which event is newer, which is why the events carry a version number. The receiver stores the version alongside the state and only applies an event when its version is higher than the stored one. This is the guard from the store snippet earlier in the post, and it is the entire fix. The test recreates the scenario by sending the events in reverse. The update, which carries version 2, goes first, and the creation event, which carries version 1, arrives afterwards:

                
const updated = buildEvent({ id: 'evt_2', type: 'subscription.updated', status: 'past_due', version: 2 });

await deliver(updated); // v2 arrives first
await deliver(created); // v1 arrives late

expect(receiver.store.subscriptions.get('sub_1')).toEqual({ status: 'past_due', version: 2 });
                

The assertion is on the final state, not on the responses. Both deliveries are accepted, and that is correct, because neither is a duplicate and neither is invalid. What matters is that the late arrival of version 1 did not overwrite version 2. I confirmed this by replacing the version check with an unconditional write, and the test failed.

There is a design decision hiding in this test, and it is worth making explicit. In our receiver the version guard protects the stored state, but the late creation event still triggers its welcome email, because the email is recorded for every creation event that is processed. Whether that is right depends on the product. Sending a welcome message to a customer who was created a minute ago and has since been updated is probably fine. Sending a cancellation notice for a subscription that was already reactivated would not be. Ordering protection is easy to apply to data and easy to forget when it comes to side effects, so a test like this one should state which of the two it is checking.

Reversing two events is a small test, but it checks a behavior that almost never fails in a developer's environment, where events are delivered one at a time in the order they were sent. It only breaks when a retry or a parallel delivery changes the sequence, and a test that creates those conditions on purpose is the only reliable way to find out. That leaves the last scenario, the one where the receiver itself is what fails.

Test 5: Retry Without Double-Processing

So far the sender has been the unreliable party. The last scenario flips that, because sometimes the receiver is the one that breaks. The database is down for a moment, a dependency times out, or the handler throws. The sender's retry logic exists for exactly this, but it only works if the receiver tells the truth about what happened. A failure has to look like a failure, and a retry has to be treated as a fresh attempt.

The receiver handles this in the last step of the request. The processing is wrapped in a try block, and if it throws, two things happen. The response is a 500, which tells the sender the event was not applied and should be sent again. And the event ID is removed from the claimed set, which is the part that is easy to forget:

                
try {
  await store.apply(event);
} catch {
  claimed.delete(event.id);
  return reply(res, 500, { error: 'processing failed' });
}
                

Releasing the claim matters because of the duplicate check from the previous tests. The ID was claimed before the work started, so if it were left in the set after a failure, the retry would be recognized as a duplicate and answered with a friendly 200. The sender would then consider the event delivered, and an event that was never applied would be lost with no error anywhere. The two protections that make the receiver safe, ignoring duplicates and accepting retries, pull in opposite directions, and this one line is where they are reconciled.

To test it, the receiver has to fail on demand. The store has a small hook for that. Calling failNext sets a counter, and the next call to process an event throws when the counter is above zero. Because it is a test double inside the store, the receiver's real error handling is exercised without needing a real outage:

                
if (failuresLeft > 0) {
  failuresLeft--;
  throw new Error('database unavailable');
}
                

The test then walks through the whole life of the event. The first delivery must fail with a 500, and no email may exist yet. The retry of the same event must succeed and be reported as processed rather than as a duplicate. And a third delivery, standing in for any further retry the sender might make, must now be treated as a duplicate, with exactly one email in total:

                
receiver.store.failNext(1);

const first = await deliver(created);
expect(first.status()).toBe(500);
expect(receiver.store.emailsSent).toHaveLength(0);

const retry = await deliver(created);
expect(retry.status()).toBe(200);
expect(await retry.json()).toEqual({ status: 'processed' });

const afterRetry = await deliver(created);
expect(await afterRetry.json()).toEqual({ status: 'duplicate' });
expect(receiver.store.emailsSent).toHaveLength(1);
                

Each of the three steps rules out a different bug. A receiver that answers a failure with a 200 fails the first. One that forgets to release the claim fails the second, because the retry comes back as a duplicate. One that does release the claim but then fails to record it again after the retry fails the third. I removed the line that releases the claim and ran the suite, and this test was the only one that failed.

There is one limit to be honest about. In this example the failure happens before the store changes anything, so a retry starts from a clean slate. Real failures can be messier. A handler that writes to the database and then crashes before sending the email has done half the work, and a retry has to finish the job without repeating the half that already happened. That needs the side effects themselves to be idempotent, and it is a harder test to write than anything here. What this one proves is the contract between the receiver and the sender: failures are reported honestly, and retries are allowed to succeed.

Proving the Tests Can Fail

Eight green tests are a comforting sight, but green only means that nothing went wrong. It does not mean the tests would have noticed if something did. A test that cannot fail is worse than no test, because it keeps telling you everything is fine. The only way to find out is to break the receiver on purpose and see whether the suite objects, which is the idea behind mutation testing.

A tool can generate these changes automatically, but for a receiver this small it is just as easy to do by hand. Each mutation below is a single temporary edit to the receiver or the store, followed by a run of the suite. Here is the one that is hardest to catch. It moves the claim from before the processing step to after it, so the code still looks reasonable and still works for any request that arrives alone:

                
await store.apply(event);
claimed.add(event.id); // was before the await
                

I applied seven mutations in total, one at a time, restoring the original code after each. The table shows which tests turned red:

Mutation Tests that failed
Skip the signature check Tampered body, wrong secret
Skip the timestamp check Stale timestamp (replay)
Remove the duplicate check entirely Redelivery, simultaneous deliveries, retry
Claim the event ID after processing Simultaneous deliveries only
Write state without the version check Out-of-order events
Keep the claim after a failure Failure and retry
Return 200 when processing fails Failure and retry

Every mutation turned at least one test red, and none of them failed more than three, so the suite is both sensitive and specific. The results also line up with the design. Each check in the receiver is guarded by a test that targets it, and the failures point to the broken behavior instead of setting off a wall of unrelated red.

Two rows are worth a closer look. Removing the duplicate check entirely fails three tests, which is the loud outcome. Moving the claim after the processing step fails exactly one, the simultaneous test, while the sequential duplicate test stays green. That is the gap described earlier, and the table is the evidence for why that test is in the suite. Similarly, "Keep the claim after a failure" and "Return 200 when processing fails" are different bugs, but the same test catches both. That is acceptable here because they break the same contract, and it is also a reminder that the table shows where tests overlap as well as where they are strong.

Limits of This Approach

This suite tests a receiver that lives in the same process as the tests, which is what makes it fast and lets it assert on the store directly. Against a deployed receiver that shortcut is not available. The same requests still work, but the outcome has to be observed through the application itself, such as an API that lists subscriptions or a test-only endpoint that exposes the email log.

The receiver's memory is the other boundary. A set of event IDs in a single process disappears on restart and is invisible to a second instance, so a duplicate could slip through after a deploy or a scale-out. In production the guarantee has to come from shared storage, typically a unique constraint on the event ID, and the concurrency test should then be run against that real database.

Finally, our sender is a simulation. The header names, the signed string, and the retry schedule here are a simplified version of what real providers do, and each provider has its own. The tests prove the receiver's logic, not its compatibility with a specific provider, so the signature scheme should always be checked against that provider's documentation.

Conclusion

A webhook receiver does not control when events arrive, how often, or in what order, so its tests cannot rely on a single well-behaved request. Playwright's request fixture makes it easy to act as a badly behaved sender: forging a signature, replaying an old request, sending the same event five times at once, reversing two events, and failing a delivery on purpose. Each of those is a short test, and each one guards against a way of charging a customer twice. Just as important, breaking the receiver deliberately showed that every test fails when it should.

The full code example from this post, including the receiver, the fixtures, and all eight tests, is available on our GitHub repository.