Selected workCircuit 3 · Tasks
Task platform
A new platform designed for tens of millions of requests a day
- My role
- Architect and main engineer · acting team lead
- When
- ABM Egypt · Nov 2025 – present
- Where it stands
- Running · customers move next
What I built
The successor to the platform every customer uses today: the API, admission, queues, autoscaling, deploys and monitoring.
Why it mattered
The current platform carries the company. Its successor has to take far more load, never lose or double-finish a task, keep cloud spend under control, and let customers move without changing a line of their code.
Illustration running in your browser. It follows the real system’s rules in simplified form; names and numbers are invented, and nothing here touches the real system.
What made it difficult
Races between servers
Four API servers accept work at the same time, and none of them may admit a task twice.
Failures in the middle
Handing a task to the queue can fail after it was already accepted. That step has to undo itself cleanly.
Scaling costs money
More servers means more spend. The autoscaler needs a hard ceiling, and it has to explain its choices.
How I solved it
Accept work in one atomic step
Small Lua scripts on Redis admit each task with per-client limits and idempotency keys, and roll the step back if publishing fails.
Queues written as code
52 queues cover 14 task types, each with live, test and qualification lanes, plus dead-letter and retry queues. Changes go through review.
An autoscaler with a budget
One engine sizes all 14 worker types, picks the cheapest server mix, stays under spend and server limits, and can run in shadow mode first.
Same API for customers
Compatible endpoints mean customers can move without code changes. Deploys roll out server by server and stop when a health check fails.
What I delivered
Built and deployed on its own nine-server production environment, with 47 alert rules and Grafana dashboards written as code.
Customers have not moved to it yet: the current platform still serves all of their traffic, and the edge router is the path they will move over on. 10–50 million requests a day is the design target, not measured traffic.
I own the architecture and wrote about 99% of its 1,177 commits. That includes work by AI coding agents that I directed through written specs and reviewed.
Built with
- Laravel 12
- Octane / FrankenPHP
- Redis + Lua
- RabbitMQ
- MySQL 8.4
- Caddy
- Hetzner Cloud API
- Prometheus
- Grafana