microservice-optimization
Two strategies for cutting end-to-end latency by changing the shape of the
request graph, not by tuning parameters inside a fixed topology.
Tuning within the current topology (connection pools, gateway upstreams, cache
usage, database indexes, worker counts, serialization) is covered by the
orchestrator's standing task guidance and is not repeated here. Reach for this
skill when the evidence points at the graph rather than at one service's
configuration.
Establish the evidence first
The two roles see different evidence. Neither should act on the other's.
Orchestrator. You see the profiler's prose summary, not the trace graph.
Select a strategy only when that summary names a specific bottleneck: a
service, an operation, or an edge with critical-path contribution. A ranking of
services by aggregate latency is not sufficient. Aggregate cost and
critical-path contribution are different quantities, and the expensive service
is frequently not the one holding the request open. When the summary names no
edge, plan a profiling round rather than guessing one.
Implementer. You can read the graph. Confirm the edge or the structure
before editing:
trace_graphs() to locate schema-v2 graphs.
critical_path(path=..., telemetry_path=...) for wall-clock attribution.
- Read
nodes_by_contribution for ranked contributors, and the representative
segments for the order in which calls actually occur.
Overlapping sibling calls are not additive. Two calls each showing 10ms of
inclusive latency may cost 10ms together, not 20ms, in which case
parallelizing them gains nothing. Representative segments distinguish
sequential work from overlapping work; flat span aggregates do not.
Critical-path scope is synchronous_request. Async and linked relationships
are excluded and counted in async_relationships_excluded. Work moved off the
synchronous path leaves the critical path by construction, so a before/after
critical-path comparison cannot by itself show that the work got cheaper.
Confirm every claim against the benchmark's primary_value.
Strategy 1: change when work happens
Two forms, in increasing risk.
Parallelize independent sequential calls. Two calls issued one after the
other, where the second does not consume the first's result, can be issued
concurrently. Verify the independence in the code, not from the trace: the
trace shows they are sequential, not that they must be.
Preconditions: no data dependency, no ordering requirement the correctness
contract relies on, and a downstream that tolerates the added concurrency.
Check the second condition against the accuracy oracle's properties, not
against the benchmark.
Move work off the request path. Precompute it, do it at write time, or run
it in the background. This removes the work from the measured path rather than
making it faster.
This form changes consistency semantics, and that is where it fails. The
accuracy oracle independently checks read-your-write behavior, index
invalidation, deletion, and isolation. Work deferred out of a write path can
improve the benchmark and fail accuracy, which is a rejected round, not a
tradeoff. State which property you are relying on staying true before making
the change.