Benjamin Cane
Portrait of Benjamin Cane
Benjamin Cane
July 23, 2026
reliability
a close up of a network with wires connected to it
Photo by Albert Stoynov on Unsplash

The closer to the edge, the more stable a platform must be.

The closer a component is to the customer, the greater its responsibility for keeping the entire platform available, even when everything behind it is having a bad day.

Not All Services Carry the Same Reliability Burden

Let’s consider a typical platform.

Customer -> Load Balancer -> API Gateway -> Orchestrator -> Microservices -> Database

Every layer has a different job. But every layer also has a different level of responsibility for resiliency.

As you move toward the customer, that responsibility increases.

Deep Services Focus on Business Logic

At the deepest layers of the platform, services are usually focused on business capabilities.

They process orders, transfer money, manage inventory, and store data.

These services often have databases, business rules, stateful operations, and multiple dependencies.

Resiliency matters, but it’s often focused on correctness.

If a database call fails:

  • Should the transaction roll back?
  • Should the service fail over?
  • Should a compensating transaction occur?

These services are primarily concerned with business outcomes.

The Middle Layers Absorb Failures

Move up a layer, and you often find orchestrators and workflow services. These components coordinate work across multiple services.

If one service fails, the orchestrator may retry, execute fallback logic, trigger compensating actions, or roll back a workflow. Their job is not just executing business logic, it’s ensuring execution succeeds despite failures.

The Edge Exists to Protect Everything Behind It

At the edge, things change.

Load balancers and API gateways are often stateless, dependency-light, highly available, and extremely fast.

Why?

Because their primary responsibility is availability. Everything behind them is allowed to fail, and they absorb as much of that failure as possible.

They:

  • Route around failures
  • Shed load
  • Fail over traffic
  • Enforce timeouts
  • Apply retries
  • Protect backend systems

The edge isn’t just resilient for itself. It’s resilient on behalf of everything behind it.

Final Thoughts

The deepest services in a platform should be focused on business logic. The edge should be focused on availability.

The more failures your edge can absorb, the less every downstream service needs to care. That’s why the closer you get to the customer, the more stable the platform must become.

Discuss on LinkedIn Newsletter Back to all posts

More to Read

  • July 16, 2026 Sometimes the most resilient thing a system can do isn’t retry reliability
  • July 9, 2026 Should retries and timeouts live in your application or your service mesh? reliability
  • July 2, 2026 Need to migrate from one database to another without downtime? architecture
  • June 25, 2026 Glue Services: Part Two — Data Synchronization architecture
  • June 18, 2026 When modernizing legacy systems, don’t be afraid to build glue services architecture

Practical engineering notes by Benjamin Cane.