Service Mesh definition
A service mesh is an infrastructure layer that manages communication between services in a distributed application. It uses proxies running alongside each service to handle encryption, authentication, retries, timeouts, traffic routing and telemetry, so these concerns are configured centrally instead of being coded into every service. Istio and Linkerd are well-known examples.
How does a service mesh work?
A service mesh has two parts. The data plane consists of lightweight proxies, commonly Envoy or Linkerd's own proxy, that intercept all network traffic to and from each service, traditionally as a sidecar container in every Kubernetes pod. The control plane configures those proxies with policies: which services may talk to each other, how to retry failures and where to send traffic.
Because the proxies handle the network, application code simply calls another service by name. The mesh encrypts the connection with mutual TLS, applies timeouts, balances load, and records metrics and traces. Sidecarless designs such as Istio's ambient mode, now generally available, reduce overhead by moving some proxy work from per-pod sidecars to shared node-level components.
Key features of a service mesh
Meshes bundle several capabilities that would otherwise need libraries in every language your services use. Implementing them consistently across Java, Go, Node.js and Python services is hard, which is one of the strongest arguments for moving them into the infrastructure layer instead.
- Mutual TLS encryption and service identity between all services.
- Fine-grained authorization policies for service-to-service calls.
- Retries, timeouts and circuit breaking.
- Traffic splitting for canary releases and A/B routing.
- Golden metrics, such as latency, traffic and errors, for every service.
- Distributed tracing headers propagated across calls.
Benefits and costs
A mesh gives platform teams consistent security and observability across services written in different languages, without asking every team to implement the same logic. It supports zero trust networking inside the cluster and makes progressive delivery, such as sending a small share of traffic to a new version, much easier.
The costs are added latency per hop, extra memory and CPU for proxies, and significant operational complexity. Upgrading the mesh, debugging proxy configuration and understanding failure modes require dedicated skills. For a handful of services, the overhead is rarely justified.
Popular service mesh options
Istio is the most feature-rich option, built on Envoy and offered as a managed add-on by several cloud providers. Linkerd, a graduated Cloud Native Computing Foundation project, emphasizes simplicity and low resource use. Consul spans Kubernetes and virtual machines, and Cilium provides some mesh features with eBPF in the Linux kernel, avoiding sidecars for parts of the work.
Whichever you choose, run it first in a staging cluster with realistic traffic, measure latency and resource overhead, and practice upgrades before relying on it in production. A mesh that the team cannot confidently upgrade becomes a risk of its own.
When should you use a service mesh?
Consider a mesh when you run many services on Kubernetes, need encrypted and authorized traffic between them for compliance, and want uniform observability and traffic control. Start smaller with a library approach or your cloud provider's managed mesh if needs are modest. Nexzem's DevOps engineers usually introduce a mesh only after a platform has outgrown simpler options such as ingress controllers and network policies.