Why Your First Distributed System Should Be a Task Queue (Not a Microservice)

The Problem That Forces Your Hand

Your web application just hit a wall. Users are uploading images, and your server blocks for thirty seconds while ImageMagick churns through each resize operation. The browser times out. Users refresh. Your server dies under the load of duplicate requests.

This is where most developers first encounter the need for distributed systems. Not because they read about microservices on Medium, but because their monolith can no longer handle work that takes longer than an HTTP request timeout. The solution isn’t to split everything into services. It’s to move that slow work somewhere else entirely.

Event-Driven Architecture: Your First Pattern

Event-driven architecture sounds complex, but it starts simple. When a user uploads an image, your web server doesn’t process it immediately. Instead, it writes a message to a queue saying “resize image X to these dimensions” and returns a 202 Accepted status. A separate worker process picks up that message and does the actual work.

Redis with its LIST operations makes an excellent first queue. Push jobs with LPUSH, pop them with BRPOP. Your web server stays responsive because it never blocks on slow operations. Workers can crash and restart without losing data because Redis persists the queue to disk. You’ve just built your first distributed system with about twenty lines of code.

The pattern works beyond image processing too. Email sending, PDF generation, data exports, webhook deliveries. Any operation that takes more than a few hundred milliseconds becomes a candidate for async processing. Start with one queue for one type of job. I know it’s tempting, but resist the urge to build a generic job system until you understand the specific patterns of your workload.

Data Consistency: Where Things Get Interesting

The moment you have workers processing jobs in the background, you face the dual-write problem. Your web server needs to save the uploaded file record to the database AND publish the resize job to the queue. What happens if the database write succeeds but the queue publish fails? You have an image record with no resized versions.

The transactional outbox pattern solves this problem. Instead of writing directly to the queue, write the job data to a jobs table in the same database transaction as your main record. A separate process polls this table and publishes jobs to the actual queue, marking them as published. If anything fails, you can retry or alert on stuck jobs.

PostgreSQL’s NOTIFY/LISTEN makes this pattern particularly clean. Your job publisher listens for notifications on new rows, avoiding constant polling. The database becomes your source of truth, and the queue becomes an optimization for work distribution. This pattern scales to millions of jobs without exotic infrastructure.

Service Decomposition: When and How to Split

The pressure to split into microservices usually comes from team growth, not technical necessity. When your codebase has distinct domains owned by different teams, service boundaries start making organizational sense. The billing team shouldn’t have to worry about deploying when the recommendations team pushes changes.

But here’s what I learned the hard way: start with a shared database. Yes, this violates microservice orthodoxy. But database transactions solve consistency problems that distributed systems make exponentially harder. When the user service and the billing service both need to update records atomically, a database transaction is your friend. Distributed transactions are definitely not.

Extract services by vertical slices, not horizontal layers. Don’t create separate services for “user management” and “user data.” Create a complete user service that handles authentication, profiles, and preferences. Give each service a clear business purpose and the data it needs to work independently. The goal is reducing chatter between services, not maximizing service count.

Message Passing: Beyond Simple Queues

As your system grows, point-to-point queues become limiting. When user registration needs to trigger email sending, analytics tracking, and recommendation engine updates, you don’t want the user service knowing about all these downstream systems. This is where publish-subscribe messaging really helps.

Apache Kafka works well here because it treats messages as immutable events in an append-only log. When a user registers, publish a “UserRegistered” event. The email service, analytics service, and recommendation service each maintain their own offset into this event stream. They can process events at their own pace, replay missed events, or spin up entirely new services that consume historical data.

The key insight is designing events as facts about what happened, not commands about what should happen. “UserRegistered” with user details is better than “SendWelcomeEmail” with an email template. Facts allow new consumers to extract value you didn’t anticipate when you first published the event. Commands lock you into specific workflows.

Starting Your Journey

Build your first distributed system to solve a real performance problem, not an imaginary scaling problem. Start with a task queue backed by Redis or PostgreSQL. Add the transactional outbox pattern when you need stronger consistency guarantees. Extract services when team boundaries make coordination harder than technical integration.

The patterns I’ve described here handle most real-world distributed system needs. They’re battle-tested, well-understood, and debuggable with standard tools. Resist the temptation to jump straight to orchestration platforms or service meshes. Master these fundamentals first. Your future self, debugging a production incident at 2 AM, will thank you for choosing boring, reliable patterns over the flashy complex ones.