Skip to content
All blog posts

Concept and method · Blog

Load Test - Did your load test actually generate the load?

Seeing 1,200 virtual users on a dashboard does not show that a high request rate was produced; think time, iteration duration and dropped iterations change the result directly. This article sets out how an HTTP API load test with k6 is sized, run and read, in the order that matters: prove the load generator first, then interpret the system.

The most common mistake in load testing is confusing seeing a number on a dashboard with proving that the load behind it was really produced. The sentence “1,200 VUs ran” says nothing on its own: those VUs may be producing only 30 requests per second because of think time; dropped_iterations may show that the target arrival rate was never reached; or the bottleneck may sit in the load generator’s CPU rather than in the system under test. In all three cases, the sentence “the system handled 1,200 users” in the report is wrong.

This playbook has one guiding principle, and it governs every step of interpreting a result: prove the load generator first, then interpret the system. The existence of a metric does not show that the load behind it is real, just as a control being defined does not show that it is enforced. What follows sets out how to scope, size, run and - most importantly - validly interpret HTTP API load tests with k6. Red in Pulse also covers two methods for source IP diversity: binding IPv6 addresses assigned to the server to k6 VUs for small and medium tests, and using an API Gateway endpoint pool where broader IP and geographic path diversity is needed.

All examples are for test environments authorized in writing only; they contain no real target, IP, domain or credential.

Note: the common language of load testing is English; terms such as VU, RPS, latency and throughput are used as they are because they are industry standard, but each is explained below with an example. Do not move on to the formulas before reading that section - every calculation that follows is built on these concepts.


Before you start: core concepts

Session. A single visit by a user, from entering the system to finishing their work and leaving. If the same person enters and leaves three times during the day, that is three sessions. Example: a user opened the application, browsed three screens, submitted a form and closed it - that is one four-minute session.

Peak hour. The 60-minute window with the highest traffic. Capacity is determined by this peak, not by the daily average; size against the average and the system collapses at the peak. “Peak hour session count is 18,000” means: 18,000 visits arrived at the system during the busiest hour.

VU - Virtual User. The concurrent “virtual user” k6 runs to imitate a real one. 500 VUs means 500 virtual users running the scenario over and over at the same time. Example: if there are on average 1,200 active sessions in the system at once during the peak hour, our target is 1,200 VUs imitating them.

Iteration. One VU running the whole scenario once, start to finish. If your scenario is “view the list → open the detail → submit the form”, one pass through those three steps is one iteration. An iteration usually produces more than one HTTP request - this distinction is critical (you will see it under RPS below).

Think time. A real user does not press buttons without pause; they read, type, hesitate. The sleep() we put in the scenario imitates that human pause. Example: if a user studies a page for 3 seconds and then clicks, roughly 3 seconds of think time goes into the scenario. Forgetting think time creates an intensity that does not exist in reality and inflates the result.

RPS - Requests Per Second / Throughput. How many requests the system handles per unit of time. This is the “real intensity” measure of a load test. The most frequent mistake is confusing VUs with RPS - they are not the same thing:

Example: 1,200 VUs, sending 6 requests per iteration, one iteration taking 240 seconds.
Expected RPS = 1,200 × 6 / 240 = 30 RPS

So seeing 1,200 VUs does not prove that a high request rate is being produced; think time and iteration duration change the result directly. “How many users” and “how many requests per second” are targeted separately and verified separately.

Latency. The time from the moment a request is sent to the moment the response arrives; this is the “speed” the user feels. It is not a single number but a distribution - which is why we look at percentiles rather than the average. p95 = 750 ms means: 95% of requests completed in under 750 ms, and 5% were slower.

Tail latency. The requests at the slowest end of the distribution (p95, p99, p99.9). An example explains why they matter:

990 of 1,000 requests : ~100 ms
The remaining 10      : ~5,000 ms
Average ≈ 149 ms      → looks "perfectly fast"
p99     ≈ 5,000 ms    → reality: one user in every hundred waited 5 seconds

The average describes the system on a good day, the tail describes it at its worst - and user satisfaction is lost at the worst moments. So the decision is made by reading p95 and p99 together, not avg.

Check and threshold. These get confused, but they do different jobs. A check looks at whether a response is functionally correct (is the status 200, is there an id in the body) but does not on its own fail the test. A threshold is the rule that decides whether the test passed or failed (for example “GET p95 must be under 1000 ms”). Example: checks: rate>0.99 and http_req_duration p(95)<1000 - one audits correctness, the other a speed limit. What stops CI/CD and produces a FAIL is the threshold.

Closed and open model. Two different logics for generating load; which one you choose determines the question the test asks.

  • Closed model (ramping-vus): the VU count is fixed. If the system slows, iterations take longer and RPS falls. Question: “Does it handle this many concurrent users?” Analogy: a fixed number of customers, where a new one enters only when another finishes.
  • Open model (constant-arrival-rate): the arrival rate is fixed. Even if the system slows, the same number of requests keeps arriving per second; the VUs needed to sustain that rise. Question: “Does it handle this many requests per second?” Analogy: customers entering the door at a constant rate, with entry not slowing down even as the room fills.

Use the closed model to verify a real user count, and the open model to hold a fixed throughput/RPS target.

Ramp-up / plateau / ramp-down. The three phases of a run: the ramp-up where you raise load gradually, the plateau where you hold it at the target, and the ramp-down where you lower it slowly. The SLA decision is made from the plateau window only, because cold start during the ramp and the descent phase both distort the average.

dropped_iterations. In the open model, iterations that could not be started because the system (or the load generator) could not keep up with the target arrival rate. Above zero it means the target load was never produced, and in that case the result cannot be used as a capacity claim.

These twelve concepts are the glossary for the rest of the playbook. The formulas and tables that follow assume you know these terms.


1. Inputs and outputs of the playbook

Five inputs must be ready before a run starts: a completed scope form, an approved user/RPS profile, the endpoint and GET/POST distribution matrix, thresholds and stop criteria, and a separate run card for each run.

Four outputs must be produced when the run ends: the k6 aggregation summary, timestamped raw metric output, correlation with application and infrastructure metrics, and a comparison report carrying a PASS/FAIL/PARTIAL/BLOCKED verdict.

Everything between those two - sizing, running, interpreting - serves the same discipline: a result becomes a capacity statement only after the load is proven to have actually been produced.


2. Sizing: where do the numbers come from?

100, 1,000 and 5,000 are test steps; they are not assumed to be the system’s real nominal load. The real target is calculated where possible from the busiest hour in production telemetry or RUM data. The daily average is not used as a capacity target because it hides short-lived peaks.

2.1 Nominal VUs: the load-testing form of Little’s Law

Concurrent users = arrival rate × time spent in the system

Nominal VUs = peak-hour session count × average session duration (s) / 3,600

Example: 18,000 sessions/hour × 240 s / 3,600 = 1,200 VUs. This value represents only the current peak load. Growth and safety margin are applied separately and visibly:

Planned VUs = nominal VUs × growth factor × safety factor
Example: 1,200 × 1.25 × 1.20 = 1,800 VUs

The factors are not buried silently in the result. The report keeps the distinction between 1,200 current nominal, 1,500 growth target and 1,800 safe verification. Where no data exists, the 100/1,000/5,000 levels are run as exploration steps and are not claimed to be production capacity.

2.2 Calculating VUs from volume

If the peak-hour page or transaction volume is known instead of the session count:

Hourly iterations = peak-hour transaction count / transactions per iteration
VUs = hourly iterations × average iteration duration (s) / 3,600

Example: 100,000 pages/hour, 5 pages per iteration, average 135 s → 20,000 iterations/hour → 750 VUs. Critical warning: “page” is not a synonym for an API request. Opening one screen can produce several API calls; the real number of HTTP requests per iteration is measured separately in the k6 scenario.

2.3 VUs, iteration rate and RPS are not the same thing

This is the most frequently skipped point in the playbook:

Expected RPS  = VUs × requests per iteration / average iteration duration (s)
Required VUs  = target RPS × average iteration duration / requests per iteration
Iteration rate = target RPS / requests per iteration

1,200 VUs sending 6 requests inside a 240-second iteration produce an average of 30 RPS. So the sentence “we saw 1,200 VUs” proves nothing about throughput; think time and pacing change the result directly.

The difference between the closed model (ramping-vus) and the open model (constant-arrival-rate) follows from the same point: in the closed model VUs are fixed and RPS falls as the system slows; in the open model the arrival rate is fixed and the VUs needed rise as the system slows. An initial estimate for an arrival-rate test:

Required active VUs ≈ iteration rate × p95 iteration duration (s)

preAllocatedVUs is chosen above this estimate with a measured safety margin; maxVUs is the upper bound that prevents uncontrolled resource consumption. If dropped_iterations occurs during the run, the result is not accepted as capacity without explicitly stating that the target arrival rate could not be produced.

2.4 Sizing across multiple workflows

Total VUs are not distributed directly across endpoint percentages. Real workflows are defined first; each one’s duration, think time, request count and traffic share are recorded.

Workflow Traffic share Avg. duration Requests/iter VUs/rate
Read-only browsing
Search/listing
Sign-in
Create/update data
Long operation/polling

In the closed model the first distribution is made with workflow VUs = total VUs × traffic share; because request distribution can drift across flows with different durations, the resulting request-tag ratios are measured and the scenario VUs are recalibrated. In the open model, giving each workflow its own arrival-rate scenario is usually more controlled.

2.5 Single-use data and account pools

If every VU uses a unique account, preparing data only for the peak VU count is not enough - ramp, plateau and repeated iterations all count:

Iteration count = VUs × window duration / average iteration duration
Data rows = iteration count × rows consumed per iteration

On a linear 0 → peak VU ramp the average load is ≈ peak VU / 2; the VU × duration area of each ramp step is summed separately. Extra margins are added to the pool: 20-25% reserve for uneven distribution across load generators, an error reserve for retries and early termination, separate data for setup/teardown/smoke, and a clean separate set for reruns. Data requiring uniqueness is partitioned across generators in advance, not consumed from a shared file in a race at run time.

2.6 Bandwidth: the load generator must not be the bottleneck

Average Mbps = RPS × (avg. request bytes + avg. response bytes) × 8 / 1,000,000
Total GB = total iterations × transfer bytes per iteration / 1,000,000,000
Total VU-minutes = (peak VU × ramp min / 2) + (peak VU × plateau min)

The calculation must show upload and download separately, verify protocol/header/TLS overhead by measurement, and treat edge-origin and cross-region transfer separately where present. Compression, cache hit ratio and error response size all change the total cost.

2.7 Calculation work card

Filled in before every main run:

Data window and timezone          :
Peak-hour session count           :
Average / p95 session duration    :
Workflows and traffic shares      :
Requests per iteration            :
Average / p95 iteration duration  :
Calculated nominal VUs            :
Growth / safety factor            :
Planned VUs                       :
Target average / peak RPS         :
Required test accounts / data     :
Estimated upload / download       :
Selected executor                 :
Assumptions and data gaps         :

Calculator output does not replace production telemetry; it only feeds the test plan with explicit assumptions. (Reference methodology: the peak-hour/session-duration, test case size and bandwidth models in the Web Performance load test calculator, cross-checked against SmartBear’s publicly available Load Testing 101 material.)


3. Load levels, durations and the progression rule

Levels are not applied back to back in a single run; each level is run separately and ramped gradually within itself. These are verification steps and do not replace the calculated capacity from Section 2. If the calculated target falls between two steps, a dedicated profile is created for that value.

Level Ramp-up Steady load Ramp-down Total
100 2 min 10 min 3 min 15 min
1,000 5 min 15 min 5 min 25 min
5,000 10 min 20 min 5 min 35 min

Progression rule: Smoke PASS → 100 PASS → 1,000 PASS → 5,000. If a level fails, you do not move up to the next one. Wait between runs: ≥5 min after 100, ≥10 min after 1,000, ≥15 min after 5,000 - and if the system has not returned to its baseline, no new run starts even after the wait has elapsed.

The k6 stage definition embeds the levels in code, so they cannot be overridden by accident with --vus/--duration:

const loadProfiles = {
  '100': [
    { duration: '30s', target: 25 }, { duration: '30s', target: 50 },
    { duration: '30s', target: 75 }, { duration: '30s', target: 100 },
    { duration: '10m', target: 100 }, { duration: '3m', target: 0 },
  ],
  '1000': [
    { duration: '1m', target: 100 }, { duration: '1m', target: 250 },
    { duration: '1m', target: 500 }, { duration: '1m', target: 750 },
    { duration: '1m', target: 1000 }, { duration: '15m', target: 1000 },
    { duration: '5m', target: 0 },
  ],
  '5000': [
    { duration: '2m', target: 500 }, { duration: '2m', target: 1000 },
    { duration: '2m', target: 2000 }, { duration: '2m', target: 3500 },
    { duration: '2m', target: 5000 }, { duration: '20m', target: 5000 },
    { duration: '5m', target: 0 },
  ],
};

4. Choosing an executor: what do you want to prove?

The executor is not a style preference; it determines the question the test asks.

ramping-vus verifies the concurrent user count. As response time rises, iterations lengthen and RPS can fall; the result cannot be interpreted through VUs alone.

api_users: { executor: 'ramping-vus', stages: loadProfiles[__ENV.LOAD_LEVEL || '100'], gracefulRampDown: '30s' }

constant-arrival-rate holds a given iteration/request arrival rate. If one iteration produces several requests, rate is not RPS directly; you calculate iteration rate = target RPS / requests per iteration. If dropped_iterations > 0, the target arrival rate was not sustained; insufficient preAllocatedVUs corrupts the result from the load generator side. Because arrival rate does its own pacing, no artificial sleep() is added at the end of an iteration.

throughput: { executor: 'constant-arrival-rate', rate: 500, timeUnit: '1s',
  duration: '15m', preAllocatedVUs: 800, maxVUs: 1200 }
Goal Executor
100/1,000/5,000 concurrent users ramping-vus
Fixed users constant-vus
Fixed iteration/RPS target constant-arrival-rate
Rising throughput/stress ramping-arrival-rate
A set number of iterations per user per-vu-iterations

5. GET/POST distribution and the endpoint matrix

The traffic profile must represent real user behavior. The main profile (75% GET / 25% POST) and the write-heavy profile (40% GET / 60% POST) are defined separately.

ID Method Operation group Weight
GET-01 GET Health/readiness 5%
GET-02 GET List/search 25%
GET-03 GET Detail 20%
GET-04 GET User/account status 15%
GET-05 GET History/latest state 10%
POST-01 POST Create/submit 15%
POST-02 POST Update/cancel 7%
POST-03 POST Batch operation 3%

The write-heavy profile runs first with 100 and then 1,000 users; a 5,000 write test is performed only after data cleanup, idempotency, queue and database capacity have been verified. For every endpoint, the path, auth, request/response size, cache, idempotency and cleanup are written into a matrix, because the sizing and cost calculations feed off those values.


6. k6 scenario template

The template below shows request-type thresholds, low-cardinality tags and custom metric discipline together. The endpoints and payloads are examples; no test is run without replacing them against the current API contract.

import http from 'k6/http';
import exec from 'k6/execution';
import { check, sleep } from 'k6';
import { Counter, Rate, Trend } from 'k6/metrics';

const BASE_URL = __ENV.BASE_URL;
const RUN_ID = __ENV.RUN_ID || 'local-run';
const LOAD_LEVEL = __ENV.LOAD_LEVEL || '100';
const TRAFFIC_PROFILE = __ENV.TRAFFIC_PROFILE || 'read-heavy';
const PATH_MODE = __ENV.PATH_MODE || 'direct-ipv4';
if (!BASE_URL) throw new Error('BASE_URL is required');

const status429 = new Counter('status_429');
const businessErrors = new Rate('business_error_rate');
const getDuration = new Trend('business_get_duration', true);
const postDuration = new Trend('business_post_duration', true);

export const options = {
  scenarios: {
    api_users: {
      executor: 'ramping-vus',
      stages: loadProfiles[LOAD_LEVEL],
      gracefulRampDown: '30s',
      tags: { run_id: RUN_ID, path_mode: PATH_MODE },
    },
  },
  thresholds: {
    checks: ['rate>0.99'],
    http_req_failed: [{ threshold: 'rate<0.01', abortOnFail: true, delayAbortEval: '2m' }],
    'http_req_duration{request_type:GET}': ['p(95)<1000', 'p(99)<2000'],
    'http_req_duration{request_type:POST}': ['p(95)<2000', 'p(99)<3000'],
    business_error_rate: ['rate<0.01'],
  },
  summaryTrendStats: ['avg', 'med', 'p(90)', 'p(95)', 'p(99)', 'max'],
};

function requestParams(name, requestType) {
  return {
    headers: { 'Content-Type': 'application/json', 'X-Test-Run-Id': RUN_ID },
    tags: { name, request_type: requestType, path_mode: PATH_MODE },
    timeout: '10s',
  };
}

function recordResult(res, requestType) {
  const ok = check(res, {
    [`${requestType} status successful`]: (r) => r.status >= 200 && r.status < 300,
  }, { request_type: requestType });
  if (res.status === 429) status429.add(1);
  businessErrors.add(!ok);
  (requestType === 'GET' ? getDuration : postDuration).add(res.timings.duration);
}

function runPost() {
  const iteration = exec.scenario.iterationInTest;
  const idempotencyKey = `${RUN_ID}-${exec.vu.idInTest}-${iteration}`;
  const payload = JSON.stringify({ testRunId: RUN_ID, idempotencyKey, value: 'synthetic' });
  const params = requestParams('POST_create', 'POST');
  params.headers['Idempotency-Key'] = idempotencyKey;
  const res = http.post(`${BASE_URL}/v1/resources`, payload, params);
  recordResult(res, 'POST');
  // In a real scenario the response schema and the business outcome must be verified as well.
}

function runGet() {
  const res = http.get(`${BASE_URL}/v1/resources`, requestParams('GET_list', 'GET'));
  recordResult(res, 'GET');
}

export default function () {
  const getRatio = TRAFFIC_PROFILE === 'write-heavy' ? 0.40 : 0.75;
  Math.random() < getRatio ? runGet() : runPost();
  sleep(Math.random() * 2 + 1);
}

The difference between check and threshold is critical: a check measures functional correctness but does not on its own fail the test’s exit code; a threshold is the pass/fail criterion and is what a CI/CD decision needs. Custom metrics are separated into Counter (GET/POST/429/retry totals), Rate (business error/correctness), Trend (custom latency) and Gauge (queue depth). Every request must carry at least the name, request_type, operation, path_mode, run_id tags; a dynamic ID or a full URL is never used as name - high cardinality breaks endpoint comparison.


7. Running

Because the load profile lives in code, only LOAD_LEVEL and RUN_ID change between runs:

# Smoke - verify the load generator and the scenario
k6 run --vus 5 --duration 1m -e BASE_URL="https://api.example.invalid" \
  -e RUN_ID="smoke-001" -e PATH_MODE="direct-ipv4" scenario.js

# 100 users
k6 run --summary-mode=full \
  --summary-trend-stats="avg,med,p(90),p(95),p(99),max" \
  --out json=results/lt-100-timeseries.json \
  -e BASE_URL="https://api.example.invalid" -e RUN_ID="lt-100-direct-ipv4" \
  -e LOAD_LEVEL="100" -e TRAFFIC_PROFILE="read-heavy" -e PATH_MODE="direct-ipv4" scenario.js

The terminal summary is for the overall view only; ramp, plateau and recovery analysis is done from the time series output. To also write the aggregation summary to a file, use handleSummary:

import { textSummary } from 'https://jslib.k6.io/k6-summary/0.1.0/index.js';
export function handleSummary(data) {
  const runId = __ENV.RUN_ID || 'local-run';
  return {
    stdout: textSummary(data, { indent: ' ', enableColors: true }),
    [`results/${runId}-summary.json`]: JSON.stringify(data, null, 2),
  };
}

8. Source IP: two methods, one discipline

Separately first, then concurrently. Paths are measured in order (Direct IPv4 → Direct IPv6 → API Gateway IPv4 → API Gateway IPv6 → dual-stack), each path first with 100 users; successful paths move to 1,000. Once the 1,000 level passes, production-like mixed traffic is run (where telemetry exists, the ratios are taken from it).

8.1 Direct IPv6 for small tests (--local-ips)

k6 can distribute VUs across local source IPs; but this option does not add IPs to the operating system - the addresses must already be assigned to the interface and routed. Suitable uses: a 100-user baseline, a limited 1,000 test, IPv4/IPv6 comparison, authorized verification of source-IP-based rate limit behavior, and widening source port capacity on a single generator.

sudo ip -6 addr add <ipv6-address-1>/<prefix-length> dev <interface>
ip -6 route get <target-ipv6> from <ipv6-address-1>

k6 run --local-ips="<ipv6-1>,<ipv6-2>,<ipv6-3>" \
  -e BASE_URL="https://ipv6-api.example.invalid" -e RUN_ID="lt-100-direct-ipv6" \
  -e LOAD_LEVEL="100" -e PATH_MODE="direct-ipv6" scenario.js

Distribution behavior: IPs are assigned to VUs in order; an IP repeated in the list receives more VUs; do not assume a random IP is chosen per request; fewer IPs than VUs may be used. Take care: use only the assigned subset rather than handing over an entire /64 block; preserve keep-alive (disabling it artificially inflates source port pressure); monitor worker CPU, file descriptor and socket capacity. And explicitly: source IP diversity is not used to bypass rate limits or a WAF.

8.2 An API Gateway pool where broad IP diversity is needed

Where an IPv6 block cannot be routed, where paths from several approved regions are to be measured, or where dependence on a single generator’s network path is to be reduced, an API Gateway endpoint pool may be preferred. Do not assume a one-to-one relationship between stage count and unique source IPs - the real distribution is measured from the target logs.

import exec from 'k6/execution';
import { SharedArray } from 'k6/data';

const gateways = new SharedArray('api-gateways', () =>
  open('./config/api-gateways.txt').split('\n')
    .map((l) => l.trim()).filter((l) => l && !l.startsWith('#')));
if (gateways.length === 0) throw new Error('API Gateway list is empty');

export function gatewayBaseUrl() {
  const strategy = __ENV.GATEWAY_STRATEGY || 'sticky';
  if (strategy === 'round-robin')
    return gateways[exec.scenario.iterationInTest % gateways.length];
  return gateways[(exec.vu.idInTest - 1) % gateways.length]; // sticky per VU
}

Sticky per VU is recommended for the main load test (connection and TLS reuse are preserved); changing endpoint on every request raises connection setup cost and distorts real user behavior. When a pool is used, verify that: the endpoints proxy only to the approved target and no open or public proxy is created; path, query, body and headers are preserved; auth headers are masked in logs; the run ID appears in both gateway and target logs; quota and throttling have been confirmed; the source of a 429 can be attributed to the gateway or the origin; and temporary APIs, stages and deployments are cleaned up after the test. The goal is not to bypass a security control but to measure, in a controlled way, the effect of source IP diversity on a capacity test.


9. Interpreting results: prove the generator first

This section is the heart of the playbook. Before interpreting the system’s result, prove that k6 produced the intended load.

9.1 Load generator validity check

Metric Meaning
vus / vus_max Active and allocated VUs
iterations / iteration_duration Completed workflows and duration including think time
http_reqs Number of HTTP requests produced
dropped_iterations Iterations that could not start at the arrival-rate target
data_sent/received Worker network volume

Before accepting the result, verify: whether the target concurrency was actually reached; whether the target RPS was produced; whether the realized iteration duration is close to the value used in sizing; whether the workflow and request-tag distribution held the planned percentages; whether account and data consumption matches the estimate; why data_sent/received deviates; whether worker CPU or network is saturated; whether there are any dropped_iterations; whether the vus_max limit was hit. If load generator capacity was insufficient, the result is not reported as target system capacity.

Realized values are recalculated for the plateau window:

Real average RPS    = http_reqs / measurement duration (s)
Real iteration rate = iterations / measurement duration (s)
Real requests/iteration = http_reqs / iterations

Differences between plan and reality are explained by redirects, retries, failed checks, iterations ending early, rising response times, think time or scenario distribution. If a difference cannot be explained, the run is not accepted as comparable.

9.2 Separating thresholds, checks and latency

THRESHOLDS in the terminal summary is the first decision point: green means the criterion was met, red means the test failed, and where there is no threshold no automatic good/bad judgment is made about a metric. checks must measure not just HTTP success but functional correctness - status, schema, required fields, the POST actually being created in the backend, an idempotent repeat not producing a duplicate. Receiving HTTP 200 while producing the wrong business outcome is not a performance success.

Latency is not a single number; the sub-metrics tell you where it was lost: http_req_blocked (socket/pool), connecting (TCP), tls_handshaking (TLS), waiting (TTFB, the main indicator of gateway/origin processing time), receiving (body download). If blocked is high, look at the generator pool; if connecting/tls is high, at network, DNS, IPv6 fallback and reuse; if waiting is high, at integration and origin dependencies; if receiving is high, at response size and bandwidth.

9.3 Percentiles, ramp/plateau and correlation

The decision is not made on avg - the average hides tail latency. p(95), p(99) and max are read together; max is a single outlier and does not decide capacity on its own. Because the k6 terminal summary combines the whole run into a single aggregation, SLA assessment is always made over the plateau time window; cold start during the ramp is reported separately, and recovery is measured after the ramp-down.

Every latency or error increase is compared against infrastructure metrics on the same time axis: CPU/memory, network throughput/connections, thread or event-loop saturation, DB connection/query latency, cache hit/miss, queue depth/consumer lag, autoscaling events, gateway/integration latency, WAF and rate-limit events. Without correlation you do not write “the bottleneck is the database” or “the gateway is slow”.

Errors are also classified by owner: client (timeout/DNS/socket/TLS → load and network team), gateway (429/5xx/integration timeout → platform), security (WAF 403/rate block → security), application (origin 4xx/5xx), data (duplicate/nonce/validation), dependency (DB/cache/queue). A 429 is not counted directly as an error or a success - it is assessed separately against expected throttling behavior.


10. Interim thresholds and stop criteria

Starting thresholds until official SLOs/SLAs exist:

Load GET p95 POST p95 Errors 5xx Unexpected 429
100 ≤500 ms ≤1,000 ms <0.1% <0.05% 0
1,000 ≤750 ms ≤1,500 ms <0.5% <0.1% <0.1%
5,000 ≤1,000 ms ≤2,000 ms <1% <0.5% <0.5%

For path comparison: IPv6 p95 must not be more than 10% worse than IPv4; the success rate difference must not exceed 0.2 points; the gateway’s added p95 cost must not exceed the greater of 20% or 100 ms; the endpoint distribution difference must stay under 15%; and the source of a 429 must be attributable. These thresholds are replaced by the organization’s official values.

The test is stopped when: 5xx exceeds 2% for two minutes; p95 exceeds twice the target for two minutes; data integrity problems or duplicates appear; a queue grows uncontrollably; the DB or origin reaches critical capacity; dropped_iterations rises continuously and the worker limit is confirmed; worker CPU, network or sockets are saturated; unexpected broad 403/429 appears; IPv6 produces continuous fallback or resets; a cost or quota alarm fires; traffic is seen going to an out-of-scope address.


11. Reporting and comparison

Every run is reported with a card placing target and realized values side by side, because the whole claim of the playbook lives in the consistency of those two columns:

Run ID / Test purpose / Address path / Source IP method / Executor:
Nominal VU calculation / Growth-safety factor:
Concurrent user target/realized · RPS target/realized:
Requests-iteration target/realized · Iteration duration target/realized:
Ramp / plateau / ramp-down · GET-POST target/realized:
Accounts-data planned/consumed · Upload-download estimated/realized:
GET p50/p95/p99 · POST p50/p95/p99 · HTTP/business error · 429/403/5xx · dropped:
Gateway/integration latency · Origin saturation · DB-cache-queue · autoscaling · recovery:
Source IP count · IPv4/IPv6 ratio · Distribution per IP:
Threshold verdict · Business correctness · Overall: PASS/FAIL/PARTIAL/BLOCKED · Bottleneck · Action · Retest:
Run VUs RPS Path GET p95 POST p95 Errors 429 Dropped Recovery Verdict
Direct-v4-100 100 IPv4
Gateway-v4-100 100 Gateway
Mixed-1000 1,000 Mixed
Mixed-5000 5,000 Mixed

Closing

The value of a good load test does not show in how many VUs appear on a dashboard; it shows in how solidly you proved that the load was actually produced, that the workflow distribution held, and that the result can be explained by infrastructure metrics. The most useful report is not a long list of graphs; it is the report that states clearly that the targeted load was verified, at what point and why the system saturated, which layer the bottleneck belongs to, and how a fix will be measured again. Because in the end what we assess is not the metric displayed; it is the load produced.


This playbook is the framework of the approach the Red in Pulse team follows in performance and capacity assessments: a methodology that sizes load from telemetry, proves the load generator before the result, measures source IP and edge behavior only in an authorized and controlled way, and leaves the test environment and its cost as clean as it started. If you would like the real capacity of your APIs - the evidence of load produced, not the number shown on a dashboard - measured with this rigor, you can contact us for a scoping conversation.

Concepts and abbreviations in this article

VU (Virtual User)

The virtual user k6 runs concurrently to imitate a real one. 500 VUs means 500 virtual users running the scenario over and over at the same time.

RPS (Requests Per Second)

The number of requests a system handles per unit of time, and the real measure of intensity in a load test. It is not the same as the VU count: expected RPS is the VU count times requests per iteration, divided by average iteration duration.

Tail latency

The requests at the slowest end of the latency distribution (p95, p99, p99.9). The average describes the system on a good day and the tail describes it at its worst; user satisfaction is lost at the worst moments.

Closed and open model

In the closed model the VU count is fixed and RPS falls when the system slows; in the open model the arrival rate is fixed and the VUs needed rise as the system slows. Which one is chosen defines the question the test asks.

dropped_iterations

Iterations that could never be started in the open model because the target arrival rate was not met. Above zero it means the target load was never produced, and the result cannot be used as a capacity claim.

// BLOG

Have the real capacity of your APIs measured

Schedule a scoping call for a capacity assessment that sizes load from telemetry, proves the load generator before the result, and names the layer the bottleneck belongs to.

Load Testing: Proving the Generated Load | RinP · Offensive Security