Building email subscriptions for this blog
The Azure architecture, bugs and testing behind the blog's new email subscriptions, from the first implementation to production.
I wanted people who enjoy the blog to have an easy way to keep up with new posts. There already is an RSS feed, but not everyone understands or wants to use that.
I also wanted to control every aspect of the service. I already had Azure Communication Services available, so I chose to build the service myself. That let me decide how subscriber data was handled, leave out open and click tracking, and choose when to send each notification.
The Azure setup
The blog is a Static Web App hosted in Azure. Its managed API previously handled the contact form. Subscriptions needed queue workers and scheduled jobs, so I moved to Static Web Apps Standard with a separately deployed Function App running Node.js 22 on Flex Consumption.
Azure Table Storage holds subscription and notification state. Storage Queues carry background work. Azure Communication Services sends the mail. The simplified path is:
Browser
-> Static Web Apps Standard
-> Function App HTTP endpoints
-> Azure Tables + Queues
-> Function App queue workers
-> ACS
-> Mailbox
I had considered Cosmos DB with a service bus. If the service grows I can consider a migration but for now a storage account is totally fine to save costs.
The HTTP endpoints and workers run in the same Function App. Timers clean up expired pending subscriptions and recover interrupted notification scheduling. ACS delivery reports return through Event Grid to an authenticated webhook.
I built the replacement separately from the live site's resources, with its own storage and queues. Real test mail could only go to one approved address, enforced at the sending step. The existing contact form moved first.
Function registration needed patience. Settings changes restart the host, and indexing happens asynchronously. Trigger sync could succeed while Azure's registered function list still lagged behind. I had to wait for readiness and verify that the host and management view agreed.
The bugs
Suppose a welcome email is queued, the reader unsubscribes, then joins again with the same address. A worker that only checks the address could send old work to the new subscription. An old unsubscribe link could delete it too.
Each subscription therefore gets its own identity. Tokens and queued jobs belong to that identity, and workers check that the original subscription still exists and is active. Joining again doesn't revive the old subscription.
Unsubscribe deletes the record and associated application-held recipient data. The address-free marker recording an article's release stays, so deleting subscribers can't make the article eligible for another campaign. Provider telemetry and backups have separate retention.
The notification workflow had a more direct bug. A preview request could overwrite a stored notification row without preserving its release marker. Release an article, request a preview, and the protection against another campaign could disappear.
The fix preserves stored state and rejects previews after release. Regression tests cover that sequence and a preview racing with a release. Table Storage's ETags prevent a request from overwriting a row based on an outdated version.
Queue failures needed a separate decision. If a release is saved but queue insertion fails, I've still authorised the campaign. The API reports that outcome, and a timer can pick up the saved work.
Once a provider send is attempted, the policy is stricter. I chose one application-level attempt per email, without retrying failed or ambiguous sends. A crash after claiming the attempt can lose a message. I accepted that trade-off for the blog. Scheduling recovery can resume unattempted work, but can't repeat an uncertain send.
Testing
Useful tests followed a subscription through several actions. Sign up, confirm, unsubscribe, join again, then try an old token. Run a worker between those steps. Repeat a confirmation concurrently. Delete the subscription while work is queued.
Notification tests covered simultaneous releases, repeated workers, interrupted scheduling and fixed prepared content. They also checked the audience cutoff. A reader joining after publication but before release is eligible; someone joining after release isn't added to that campaign.
The storage assumptions needed testing too. Azurite checks exercised conflicting writes, stale ETags and transaction rollback against the emulator. By the final production build, the project-wide suite had 98 passing tests with no skips, including all three emulator tests, alongside lint, type checks and the build.
Browser checks covered signup, unsubscribe and rejoin, keyboard operation, mobile layouts and both themes. An injected API failure left the typed address available. Opening an unsubscribe link did nothing until I confirmed, which matters when email software fetches links automatically.
One live test remained inconclusive. It was meant to observe a pending subscription through its real 24-hour expiry and the next hourly cleanup. The observer stopped recording before expiry, and its intended shutdown didn't complete either. Later inspection found the record gone, but couldn't establish when it had been deleted.
I accepted that missing timing evidence as a documented limitation. Deterministic expiry tests passed; the overnight observation didn't. Processing was subsequently shut down and verified. Waiting for a test that failed to record the answer still used calendar time.
What turned up in the mailbox
Live email tests used one approved address and a temporary article on the isolated site. Confirmation, welcome and notification messages arrived. The notification reached Proton, but its headers recorded a spam action. Provider success and passing authentication hadn't established normal inbox placement.
Event Grid also needed checking end to end. Early tests showed matching events alongside failed webhook deliveries. Later checks matched an authenticated Delivered report to the current subscription's send operation.
The unsubscribe button was more surprising. My test address passed through SimpleLogin before Proton. The forwarded message's unsubscribe header pointed to SimpleLogin's alias action. Its documentation says that disables the alias by default, with an option to block a sender instead. Neither would prove deletion of the blog subscription.
Native one-click unsubscribe had another complication. RFC 8058 requires DKIM coverage of both unsubscribe headers. A Microsoft ACS support response reported that ACS's fixed signing set omitted them. The forwarded message hadn't retained the original blog signature, so it couldn't verify that coverage independently.
I launched with the personalised unsubscribe link in each email and the site's unsubscribe form. Native mail-client one-click eligibility stayed outside the accepted launch scope.
Almost three weeks later
Production cutover included inspecting the build, checking domain bindings and certificates, changing DNS, and verifying the public routes and API behaviour. The temporary test article was excluded, the 16 archive posts were protected from notification release, and deployment sent no campaign. The old site remains available for rollback.
The subscription lifecycle landed on 12 September and notifications on 20 September. On 30 September, I was still changing the Function App's instance size and checking live behaviour. Production followed on 2 October.
A signup already worked much earlier. The remaining work was finding out what happened when requests overlapped, readers left while jobs were queued, or Azure accepted one step and failed the next.
The service is live now. You can subscribe here.