Skip to content

Use p95 latency by route to find slow API endpoints

Clicktag tracks 95th percentile latency by route template rather than raw URLs to isolate slow API endpoints. By grouping parameter paths, you prevent data fragmentation and measure response times within roughly 4% accuracy.

Understanding Latency Distributions: Arithmetic Mean vs. Percentiles

An average is sensitive to the mix of fast and slow requests. Percentiles add a different view: p50 describes the middle of the distribution, p95 the threshold around which 95% of observed durations fall at or below, and p99 a point further into the tail. Repeated durations mean the split need not be exact.

Read the three together. A low p50 with a much higher p95 suggests a slower subset worth investigating. It does not identify the cause: different request sizes, dependencies, cache states, or consumer behavior could all contribute.

Clicktag uses histogram-based approximations with roughly 4% accuracy. Small differences between nearby values are not a reason to declare a regression. Compare meaningful changes over comparable windows, and keep request counts beside latency measurements.

Why Route Templates Prevent Cardinality Collapse

Percentiles require an adequate sample size to yield meaningful conclusions. If an observability tool records latency against raw request paths, dynamic URL segments will fragment the underlying dataset.

For example, if your service receives requests for /users/101/profile and /users/102/profile, tracking raw paths creates two distinct rows. Across thousands of active users, traffic splinters into thousands of separate paths that each log only a few events. Calculating a p95 latency on an endpoint with five requests produces volatile, statistically unhelpful figures. Furthermore, unbounded cardinality increases memory pressure and slows down analytical queries.

Grouping requests by route template solves this issue. A pattern such as /users/:id/profile groups related calls under a single reporting identifier. Framework adapters handle route parameterization through specific mechanisms:

  • Hono, Express, and Fastify middleware read the internal route definition directly from the router match table.
  • Next.js adapters extract the route structure from matched route segment metadata.
  • Cloudflare Workers adapters pass the raw request path directly to Clicktag, and Clicktag templates the path during server-side ingestion. Clicktag replaces numeric identifiers, UUID strings, long hexadecimal hashes, and random tokens with :id, calendar dates with :date, and email addresses with :email.

In addition, scanner probes and arbitrary paths returning a 404 status that match no registered parameter route are grouped into a fallback /* bucket. This prevents malicious fuzzing and web crawlers from degrading your endpoint reporting. You can review pattern normalization rules in the Clicktag API analytics documentation.

Weighing Request Volume Against Latency Numbers

A common mistake during performance audits is sorting endpoints exclusively by the highest p95 duration. Relying solely on that ordering often directs engineering resources toward background routes with negligible real-world impact.

Consider an internal reporting endpoint that runs once per day to export accounting files and completes in 2,800 milliseconds. Its p95 latency will report near 2,800 milliseconds. Conversely, an API endpoint for checkout or token verification might process 150,000 requests per hour with a p50 of 40 milliseconds and a p95 of 520 milliseconds.

Refactoring the accounting export saves several seconds once daily. In contrast, dropping the checkout endpoint p95 from 520 milliseconds to 150 milliseconds eliminates hundreds of cumulative minutes of latency for active customers every hour. Evaluate total request throughput alongside percentile latency to prioritize code and database updates effectively.

Comparing Endpoints to Prioritize Optimization

Before refactoring your application, inspect baseline performance metrics across your core routes over a rolling window, such as the preceding 7 days. Compare request count, server error rates, p50, and p95.

The following table illustrates a hypothetical operational breakdown across four sample routes. These values are illustrative and do not reflect specific benchmark data.

Route Template Requests (7d) 5xx Error Rate p50 Latency p95 Latency p99 Latency Priority Assessment
GET /api/v1/products/:id 650,000 0.02% 15 ms 42 ms 95 ms Low (consistent, tight tail)
POST /api/v1/orders 110,000 0.92% 75 ms 580 ms 1,200 ms Urgent (high volume, wide tail)
GET /api/v1/search 85,000 0.15% 38 ms 390 ms 820 ms High (variable query execution)
POST /api/v1/exports/archive 60 0.00% 1,900 ms 2,700 ms 2,850 ms Low (infrequent background job)

In this illustrative scenario, POST /api/v1/orders represents the highest return on engineering effort. The substantial gap between its p50 (75 ms) and p95 (580 ms) indicates that tail requests are experiencing database row lock contention or awaiting responses from downstream payment gateways. Conversely, GET /api/v1/products/:id maintains healthy consistency across all percentiles.

Collecting Route Metrics with Application Middleware

Capturing percentile metrics does not require complex tracing agents or intrusive instrumentation. The official totallytics middleware package aggregates counters and latency histograms locally in application memory, then transmits compact summaries in background batches.

In Node.js, Bun, and Deno environments, the middleware collects metrics and flushes batches every 10 seconds by default. On Cloudflare Workers, the adapter flushes metrics approximately 5 seconds after the initial incoming request using background execution hooks.

To install the middleware, add the package to your application:

npm install totallytics

Set your site API key in the environment using TOTALLYTICS_API_KEY. You can copy this value from your site configuration in the Clicktag dashboard.

The following example demonstrates registering the middleware on a Hono application:

import { Hono } from "hono";
import { totallytics } from "totallytics/hono";

const app = new Hono();

app.use("*", totallytics());

app.get("/users/:id", (c) => {
  return c.json({ id: c.req.param("id") });
});

export default app;

The middleware records HTTP methods, status codes, route templates, and execution durations. It never reads request payloads, response bodies, URL query parameters, or IP addresses. For custom instrumentation requirements and JSON submission schemas, see the Clicktag API requests specification.

Frequently Asked Questions

Why does Clicktag estimate percentiles instead of storing every raw latency value?

Writing every individual request timestamp and duration into a database causes unnecessary resource overhead at high scale. By binning request durations into log-scale histogram buckets directly inside process memory, the middleware sends fixed-size aggregation payloads. Clicktag calculates percentiles from these buckets to within approximately 4% of actual values above 1 millisecond, preserving analytical accuracy without infrastructure bloat.

How does Clicktag normalize paths in Cloudflare Workers compared to Node.js frameworks?

Node.js frameworks like Express and Hono maintain route tables in memory, so the adapter reads the matched template pattern directly from the runtime context. Because Cloudflare Workers applications often route requests through custom logic or simple fetch handlers, the Workers adapter forwards the raw path to Clicktag. During ingestion, Clicktag automatically converts UUIDs, numeric IDs, long hexadecimal strings, and random tokens into :id, dates into :date, and email addresses into :email.

What causes a wide divergence between an endpoint p50 and p95 latency?

A wide gap between p50 and p95 generally indicates intermittent resource bottlenecks rather than uniform code inefficiencies. Frequent causes include cold cache lookups, database lock queues during concurrent writes, unindexed filter parameters, external third-party service latency, or memory management pauses under load.

How does route templating handle 404 status codes and web scanners?

If a client requests a missing resource on a valid parameterized route like /users/:id and your application returns a 404, the request is logged under that route template. If an automated vulnerability scanner requests an unmapped URL that matches no application route, Clicktag aggregates the request under the catch-all /* route to keep your core endpoint lists clean.

Last updated