Home / Whitepapers / APIs and microservices
Whitepaper · APIs and microservices

APIs and Microservices Without the Distributed Monolith

Designing APIs, choosing between a modular monolith and microservices, and running integrations safely. Written for CIOs, enterprise and solution architects, engineering managers, and integration and security teams.

Reading time25 minutes
LengthApprox. 15 pages
Editionv1.0, October 2026
FormatWeb and PDF

What is inside

  • How to choose between a monolith, a modular monolith and microservices
  • API design and versioning rules that keep callers working
  • When to call an API and when to publish an event
  • Timeouts, retries, circuit breakers and idempotency, explained simply
  • Availability and event sizing formulas with a worked example
  • API security based on the OWASP API Top 10, and opening APIs to partners
  • A 36-point readiness checklist for integration and API programmes
Or start reading online ↓
Download the full PDFApprox. 15 pages, including the 36-point checklist

We confirm your email with a one-time code, then the PDF downloads straight away. No newsletters unless you ask.

01

Executive summary

Most integration problems are not about technology. They come from point-to-point connections nobody documented, services split before anyone knew where the boundaries were, and one slow system dragging others down with it. Microservices promised independent teams and faster releases. Many organisations got the opposite: a distributed monolith, where every change still touches several services, but now each call can fail over the network.

This paper is for the people who shape how systems talk to each other: CIOs, enterprise and solution architects, engineering managers, and the integration and security teams who keep it running. It sets out how to choose the right structure, design APIs that last, and make integrations safe to operate.

Five things to take away

  1. Start with a modular monolith. Split a service out only when teams block each other, or one part needs very different scale or release pace. Microservices solve an organisational problem, not a technical one.
  2. Design the contract first. Write the API in OpenAPI before the code, and never break existing callers. Add fields freely; remove them only in a new version, with notice.
  3. Use events for news, calls for questions. When several systems react to the same thing, publish an event. Long chains of synchronous calls multiply failure.
  4. Make failure small. Every call has a timeout, retries are limited and idempotent, and circuit breakers stop one slow service from taking down the rest.
  5. Know every API you expose. Most API breaches are missing checks on APIs nobody remembered. Keep a catalogue, put a gateway in front, and check object ownership in every service.
02

Why integration gets painful

Integration estates grow one project at a time. Each connection made sense when it was built; together they form a web that nobody can draw from memory. When we review integration estates, the technology varies but the problems repeat:

What we findWhat it leads to
Spaghetti connectionsEvery system talks directly to every other. Change one, and you find out what breaks in production.
Microservices too earlyA small team split one application into twenty services and got all of the complexity with none of the benefit.
Shared databasesServices read and write each other’s tables. No service can change its schema without a coordinated release.
Long synchronous chainsOne user request passes through six services. Any one of them being slow makes the whole request slow.
No timeoutsA service waits forever on a slow dependency, threads pile up, and the failure spreads to everything that calls it.
Unknown APIsOld versions and test endpoints stay open on the internet, with no owner and no monitoring.
Breaking changesA field is renamed, partners’ integrations fail overnight, and nobody knew who was calling.
The distributed monolith: services that must be deployed together, share data, and call each other in long chains. It has the coupling of a monolith and the failure modes of a distributed system. Avoiding it is the main aim of this paper.
03

Core concepts: boundaries, contracts and versions

Boundaries come from the business

A good service boundary follows a business capability with its own data and its own language: orders, payments, inventory, customers. Domain-driven design calls this a bounded context. Inside a boundary, things change together and can share a database. Across a boundary, they talk only through a published contract. If two parts of a system always change together, they belong in the same boundary, however large it looks.

Boundaries matter whether you run one deployable unit or fifty. A modular monolith uses the same boundaries as microservices, enforced as modules with separate schemas inside one application. That keeps the option to split later without paying for it now.

Contract-first API design

A good API is one a developer can use correctly from its documentation alone, and one you can change without breaking the people who depend on it.

RuleWhy it matters
Write the contract first, in OpenAPI (AsyncAPI for events)Callers and providers agree before code exists; documentation, mocks and tests come from the same file.
Model resources, not database tablesThe API can stay stable while the storage behind it changes.
One error format everywhereA clear code, message and correlation ID; callers handle errors the same way for every API.
Idempotency keys on payments and ordersA retried request never charges or ships twice.
Pagination and filters on every listNo endpoint returns everything at once, so performance stays predictable.
Validate every request against the contractBad input is rejected at the edge, not deep inside a service.

Versioning without breaking callers

Adding an optional field or a new endpoint is not a breaking change. Removing or renaming a field, changing a type, or tightening validation is. Avoid breaking changes; when one is unavoidable, publish a new major version (for example /v2/), run both for an agreed period, usually six to twelve months for partners, tell callers early, and use gateway logs to see exactly who still calls the old version before you retire it.

REST, GraphQL or gRPC

  • REST for most public and partner APIs: widely understood, easy to cache and secure.
  • GraphQL when many front-ends need different shapes of the same data, behind a well-governed schema.
  • gRPC for fast internal service-to-service calls with strict contracts.
04

Monolith, modular monolith or microservices

Microservices solve a specific problem: many teams needing to change and release parts of a system independently. If you do not have that problem, they mostly add cost. The honest comparison:

Traditional monolithModular monolithMicroservices
Best forSmall, stable applicationsOne to a few teams; a product still taking shapeMany teams; parts with very different scale or release needs
DeploymentOne unitOne unit, built from clear modulesMany units; needs a platform and automation
DataOne shared schemaOne database, a schema per moduleA database per service; harder consistency
FailureIn-process calls rarely failIn-process calls rarely failEvery network call can fail and must be handled
Team independenceLowModerateHigh, if boundaries are right
Running costLowestLowHigher: platform, monitoring, gateway, broker, people
Main riskBecomes a tangleModule rules not enforcedThe distributed monolith

Signs you should split a service out

  • Teams wait on each other to release, every week.
  • One part needs to scale very differently from the rest, such as search or tracking.
  • One part changes daily while the rest is stable.
  • A clear business boundary exists, with its own data and its own owner.

Signs you split too early

  • Every feature needs changes in several services at once.
  • Services share a database, or must be deployed together.
  • You cannot test anything without starting everything.
05

A decision framework

Three decisions shape most integration architectures. Make them per domain, not once for the whole organisation.

  1. How many deployable units? Count the teams that need to release independently. One team rarely needs more than one or two services. Split along business boundaries, one service per team at most to begin with.
  2. Call or event? If the caller needs an answer to continue, call an API. If the sender is announcing that something happened, publish an event and let others react.
  3. Change or wrap the legacy system? Put an API or adapter in front first. Rewrite only what blocks the business, one piece at a time.
SituationUseExample
Caller needs an answer to continueAPI callCheck stock before confirming an order
Several systems react to the same thingEvent“Order placed” updates warehouse, billing and notifications
The receiver may be offline for a whileEvent or queueA branch system that syncs every few minutes
Very high volume of updatesEvent streamTracking pings from thousands of vehicles
Long-running work with stepsQueue plus status APIDocument processing or loan approval

Modernising a legacy system: the strangler pattern

Put a gateway or API in front of the legacy system. Build new capabilities as separate services behind the same front. Move existing features out one at a time, routing traffic to the new version, and retire the old parts as they empty out. Callers never notice, and every step can be rolled back. Big-bang rewrites of core systems fail far more often than they succeed.

06

Reference architecture

The diagram shows a typical shape once a few domains have their own services. Callers only ever reach the gateway. Services call each other sparingly and with timeouts, share news through events, and reach legacy systems only through adapters.

CALLERSGOVERNANCE AND OPERATIONSEDGEDOMAIN SERVICESEVENTS AND ERPWeb and mobile appscustomers and staffPartnerskeys plus mTLSDeveloper portalcatalogue, docs, keysLogs, metrics, tracesOpenTelemetry, trace IDsWAFOWASP rules, bot filterAPI gatewayauth, limits, routingIdentity providerOAuth 2.0 and OIDCOrders serviceREST, own databasePayments serviceidempotency keysInventory servicegRPC, own databaseCustomer servicesyncs with ERPService databasesone per service, not sharedERP adapteranti-corruption layerEvent brokerKafka topics by domainConsumersbilling, alerts, analyticsLegacy ERPSOAP and batch filesHTTPSmTLSvalidate tokenlag, traces123456User or API trafficControl / API callData / replicationScheduled copyLogging / management

Numbered flows: (1) partners register in the developer portal and receive keys and certificates, (2) all external traffic passes the WAF, then the gateway checks the token with the identity provider, applies rate limits and routes to the right service and version, (3) orders calls payments with a timeout, a limited retry carrying an idempotency key, and a circuit breaker, (4) payments publishes “payment received” through an outbox, (5) billing, alerts and analytics react independently, (6) the legacy ERP is reached only through an adapter that translates its model.

Gateway patterns

PatternWhat it doesUse when
Edge gatewayOne entry point for authentication, rate limits, routing and loggingAlways, for anything exposed outside
Backend for front-endA thin API per channel (mobile, web) that shapes data for itChannels need very different responses
Partner gatewaySeparate keys, mTLS, quotas and a portal for partnersYou open APIs to third parties
Service meshmTLS, retries and traffic control between services, outside the codeMany services, and a platform team to run it

Keep business logic out of the gateway. It should check, limit, route and record. Gateways that start transforming data and orchestrating calls become a new monolith that only one team understands. Common choices are Kong, Apache APISIX, a cloud API gateway, or a Kubernetes Gateway API implementation.

Resilience patterns

PatternWhat it preventsRule of thumb
Timeout on every callThreads waiting forever on a slow dependencySet from the dependency’s p99 latency; always shorter than the caller’s own timeout
Retry with backoff and jitterFailing on brief network blipsTwo retries at most, only on safe errors, and only at one layer
Idempotency keyDouble charges or duplicate orders when a retry succeeds twiceClient sends a unique key; the service stores and replays the first result
Circuit breakerHammering a service that is already downOpen after a run of failures, fail fast, test again after a pause
BulkheadOne slow dependency using up all threads or connectionsSeparate pools per dependency
FallbackA whole page failing because one part didShow cached or partial data where the business accepts it

Events, done reliably

Kafka suits high-volume streams that several consumers read and may replay; RabbitMQ suits work queues and routing; NATS suits lightweight messaging. Whichever you use, five practices matter more than the product:

  • Outbox: save the change and the event in the same database transaction, then publish, so the database and the events never disagree.
  • Idempotent consumers: messages can arrive twice; processing one twice must be harmless.
  • Schemas: describe events with a schema registry (Avro or JSON Schema) and evolve them without breaking consumers.
  • Dead-letter queues: messages that keep failing go aside for a person to inspect, instead of blocking the rest.
  • Ordering: decide where order matters, and partition by the right key, such as order ID.
07

Sizing and cost

Availability of a call chain

When a request depends on several services in sequence, their availabilities multiply:

End-to-end availability = A1 × A2 × … × An
Calls reaching the deepest service = (1 + retries per layer), raised to the power of the number of retrying layers

Six services at 99.9 percent each give 0.9996 ≈ 99.4 percent, or about 52 hours of failed requests a year instead of nine. And if three layers each retry twice, one user click can become 3 × 3 × 3 = 27 calls to the deepest service, exactly when it is struggling. Retry at one layer only.

A worked example: placing an order

An order passes through orders, payments, inventory, customer and notifications, each running at 99.9 percent availability.

Design A: all synchronous callsDesign B: two calls, the rest as events
Services the user waits on52 (orders and payments)
End-to-end availability0.9995 ≈ 99.5%0.9992 ≈ 99.8%
Failed-order time per year8,760 × 0.005 ≈ 44 hours8,760 × 0.002 ≈ 18 hours
If notifications is downOrders failOrders succeed; notifications catch up later
Response timeSum of five callsSum of two calls

Sizing an event stream

Partitions ≥ peak messages per second ÷ messages per second one consumer can process
Storage per day = message rate × message size × 86,400 seconds; total = per day × retention days × replicas

A logistics company tracks 20,000 vehicles that each send a ping every 30 seconds: about 670 messages a second, or roughly 1,340 at peak. If one consumer handles 200 a second, it needs at least 1,340 ÷ 200 ≈ 7 partitions; plan 12 to 16 to allow growth, because partitions are hard to add later without reordering keys. At 1 KB per message, that is about 115 GB a day; seven days of retention with three replicas needs 115 × 7 × 3 ≈ 2.4 TB.

What microservices add to running cost

ItemPlanning range per year
Container platform for 15 to 30 services, production and non-production₹40 lakh to ₹1 crore in infrastructure
Event broker cluster (three nodes) and schema registry₹10 to ₹25 lakh
Gateway and developer portal: open source, or commercial licences₹0 to ₹60 lakh
Observability storage for logs, metrics and traces₹8 to ₹25 lakh
Platform and SRE people: 2 to 4 engineers, fully loaded₹50 lakh to ₹1.6 crore

These are planning ranges, not quotes. A well-built modular monolith for the same product might run on a few VMs and one database cluster for a small fraction of this. The extra spend is justified when it buys real team independence; if it does not, it is waste.

08

Security and governance

APIs expose business logic and data directly. The most common API breaches are not clever attacks. They are missing checks, such as letting a logged-in user read someone else’s record by changing an ID in the URL. The OWASP API Security Top 10 (2023) is a good working list:

RiskWhat it looks likeControl
Broken object-level authorisationChange /orders/1001 to /orders/1002 and see another customer’s orderCheck ownership of every object in the service, not only at the gateway
Broken authenticationWeak tokens, keys in URLs, no expiryOAuth 2.0 and OIDC, short-lived tokens, mTLS for partners
Excessive data exposureThe API returns full records and the app hides fieldsReturn only the fields the caller needs
Unrestricted resource useScraping, brute force, sudden cost spikesRate limits and quotas per caller at the gateway
Function-level authorisationA normal user can call admin endpointsSeparate admin APIs; check roles on every endpoint
Improper inventoryOld versions and test APIs still exposedAn API catalogue, and retire old versions on a schedule

Identity and transport

  • Users and apps get tokens from one identity provider using OAuth 2.0 and OIDC; tokens live minutes, not days.
  • Machine-to-machine calls use the client credentials flow with narrow scopes, never a shared admin key.
  • Partners use mTLS as well as tokens, so a leaked key alone is not enough.
  • Inside the platform, encrypt service-to-service traffic with mTLS, through a mesh or the platform’s network layer.
  • Keep secrets and keys out of code and front-end apps; load them from a vault at run time.

Opening APIs to partners

Partner APIs are a product. Give partners a portal with documentation, a sandbox with realistic test data, self-service keys, published rate limits and a status page. Agree a deprecation policy in the contract. Measure usage per partner, because that is what tells you which versions you can retire and which partners need help.

Governance that teams accept

Publish short API guidelines and enforce them automatically: lint OpenAPI files in CI, run security tests before release, and register every API in the catalogue with an owner. Log every call with caller identity and alert on unusual patterns. For regulated sectors, RBI, SEBI and IRDAI expectations on access control and audit trails, CERT-In 6-hour incident reporting and the DPDP Act all assume you know which APIs touch personal data and who called them.

09

Operating it: day 2

Observability across services

In a distributed system, a single request touches several services, so logs alone are not enough. Give every request a trace ID at the gateway and pass it through every call and event, using OpenTelemetry. Collect three views for every service: the rate of requests, the rate of errors and the duration (the RED method), and for every event consumer, its lag. When something breaks, the trace shows which hop was slow, not just that the page was.

Service levels and error budgets

Set a service level objective per API that matters to users, for example 99.9 percent of requests succeed within 300 milliseconds over 30 days. The gap to 100 percent is the error budget: about 43 minutes a month. When a team has spent it, reliability work comes before new features. This turns reliability arguments into numbers.

Release and change

  • Deploy each service independently through its own pipeline; if two must always deploy together, merge them or fix the contract.
  • Run consumer-driven contract tests (for example Pact) in CI, so a provider change that breaks a caller fails the build.
  • Release behind feature flags or with canary traffic for high-risk changes.
  • Track who calls each API version, and retire old versions on the published schedule.
  • Review dead-letter queues daily; a growing queue is a silent outage.

Ownership

Every API and every event topic has a named owning team, recorded in the catalogue. The team that builds a service runs it and carries its alerts. A small central platform or integration team owns the gateway, broker, portal and guidelines, and reviews new APIs, but does not build every integration. Central teams that build everything become the bottleneck the architecture was meant to remove.

10

A practical roadmap

Integration improvement works best in steps that each deliver something, without a big-bang rewrite.

StageTypical durationOutcome
1. Map3 to 4 weeksEvery integration, its owner, protocol and what breaks if it stops. Exposed APIs found and listed.
2. Decide the shape2 to 4 weeksDomains and boundaries, where events fit, what stays as it is, API guidelines drafted.
3. Front door4 to 8 weeksGateway, identity, rate limits, catalogue and portal in front of existing APIs. Unused endpoints closed.
4. Make failure safe4 to 6 weeksTimeouts, retries, circuit breakers and idempotency on critical paths; tracing end to end.
5. Change graduallyOngoing, by quarterWorst point-to-point links replaced first; new features as services or modules; legacy strangled piece by piece.
6. Hand over4 weeks, overlappingGuidelines, templates, review practice and runbooks owned by your teams.
11

Ten common mistakes

  1. Splitting before the boundaries are clear. The result is a distributed monolith that is harder to change than the original.
  2. Sharing a database between services. No service can change its own data without a coordinated release.
  3. Long chains of synchronous calls. Availability falls and latency adds up with every hop.
  4. Calls without timeouts. One slow dependency takes down everything upstream.
  5. Retrying at every layer. A small fault becomes a flood of traffic at the worst moment.
  6. Payments and orders without idempotency keys. A retry charges or ships twice.
  7. Breaking changes without a new version. Partners find out in production.
  8. Authorisation only at the gateway. Services trust any request that gets past it, and object-level checks are missed.
  9. Business logic in the gateway or the bus. A new central monolith that one team controls.
  10. No catalogue and no owners. Forgotten APIs stay open and unmonitored for years.
12

API and integration readiness checklist (36 points)

Use this list to score your current position. Anything you cannot tick with evidence is a gap worth closing.

Architecture and boundaries

  • Domains and boundaries mapped to business capabilities, each with an owner
  • A written reason for every service split, tied to team or scale needs
  • No database shared between services or modules
  • User-facing requests wait on no more than two or three synchronous calls
  • Events used where several systems react to the same change
  • Legacy systems reached only through adapters or the gateway
API design 6 points Resilience 6 points Security 7 points Governance and partners 5 points Operations 6 points

The remaining 30 points are in the PDF, laid out as a printable checklist.

Get the full checklist
13

Glossary

TermMeaning
Bounded contextA part of the business with its own data, rules and language; a natural service or module boundary.
Modular monolithOne deployable application built from strictly separated modules, each owning its data.
Distributed monolithServices that must change and deploy together, with the downsides of both styles.
IdempotencyDoing the same operation twice has the same effect as doing it once.
Circuit breakerA guard that stops calls to a failing dependency for a while, so callers fail fast.
Outbox patternWriting an event to the same database transaction as the change, then publishing it.
Strangler patternReplacing a legacy system gradually by routing features to new services behind one front.
mTLSMutual TLS: both sides of a connection prove their identity with certificates.
Error budgetThe amount of failure a service level objective allows over a period.

About Vakratron Systems

Vakratron Systems is a vendor-neutral infrastructure design firm. We design data centre, disaster recovery, cloud, GPU and AI platforms for enterprises and government buyers, write our assumptions down, and stay with a design until it is running and tested.