Performance and sizing
These figures come from a benchmark run on 2026-09-29 against NullWard 0.6.66-beta. Use them to size your deployment. They're relative, per-vCPU figures from a single test machine, not guarantees. See Caveats. Sizing summary With every feature turned on (OWASP CRS in block mode, ML, sign-in, rate limiting, IP intelligence, request logging and the other protections), small requests over keep-alive HTTPS: 1 vCPU 2 vCPU Maximum throughput about 490 req/s about 1,040 req/s Recommended planning figure (50% of maximum) about 250 req/s about 500 req/s Latency at 50% load (p50 / p99) 3.0 / 11.8 ms 2.8 / 18.4 ms Plan for about 250 requests per second per vCPU with everything on. At 50% of the maximum, p99 latency was 12–18 ms. Request bodies cost much more than small GETs. The OWASP Core Rule Set runs its patterns over every JSON field name and value. A benign 10 KB JSON POST manages about 6 req/s per vCPU with the CRS on. For APIs that take large bodies from trusted clients, use WAF bypass paths or whitelist rules on those paths. New TLS connections are expensive. A full TLS handshake per request is the limit for clients that don't reuse connections. Make sure any load balancer in front of NullWard keeps connections alive. Throughput roughly doubles from 1 to 2 vCPUs (2.0–2.3x with everything on). How it was measured Machine. AMD Ryzen 9 8945H, Linux under WSL2, Docker. NullWard was pinned to 1 vCPU (one hardware thread, with its SMT sibling idle) or 2 vCPUs (one full physical core, which is how cloud providers count vCPUs). The backend, tunnel connector, SurrealDB, Valkey and load generator each ran on their own cores. Validity. NullWard ran at 91–97% of its CPUs at each maximum, while the backend and load generator stayed below about 45%, so the figures measure NullWard. Every response in every run had to be a 2xx from the backend: a single error, 403 or 5xx would have invalidated the run. All 42 runs were valid. Paths. direct is client → NullWard → backend. tunnel is client → NullWard → WireGuard tunnel → connector → backend. TLS means HTTPS terminated at NullWard (TLS 1.3, ECDSA P-256) with plain HTTP to the backend. plain means HTTP throughout. Requests. GET / with a response of about 180 bytes, or a benign 10 KB JSON POST. Maximum throughput. Closed-loop saturation at 16, 32, 64 and 128 connections, stopping when throughput gained less than 3% or p99 passed 1 s. Latency. Open-loop load at 25%, 50% and 75% of that maximum, 10 s each. Feature levels. Each level adds to the one before. Response caching was off throughout. Level Adds L0 Plain reverse proxy L1 OWASP Core Rule Set, block mode L2 ML WAF, detect mode L3 Email-code sign-in, with a session cookie checked on every request L4 Rate limiting (never triggered), IP intelligence, 4xx behaviour scoring L5 Request logging L6 HSTS, body size limit, forwarding headers and a cookie rule: everything on Results Throughput over keep-alive connections (requests per second) Path Protocol Level 1 vCPU 2 vCPU direct TLS L0 plain proxy 3,802 7,863 direct TLS L1 + CRS 1,428 2,176 direct TLS L2 + ML 1,066 2,263 direct TLS L3 + sign-in 492 1,095 direct TLS L4 + protections 537 1,245 direct TLS L5 + request log 506 1,085 direct TLS L6 everything 487 1,037 direct plain L0 4,395 7,642 direct plain L1 1,233 2,308 direct plain L6 470 1,083 tunnel TLS L0 2,969 4,557\* tunnel TLS L1 750 1,883 tunnel TLS L6 429 948 tunnel plain L0 3,502 4,390\* tunnel plain L1 1,217 2,157 tunnel plain L6 504 1,014 direct TLS L0, 10 KB JSON POST 3,196 4,600 direct TLS L1, 10 KB JSON POST 6 12 direct TLS L6, 10 KB JSON POST 6 14 \* The tunnel connector was near its own CPU limit at these points, so they may understate NullWard's capacity. New TLS connections (one full handshake per request) Level 1 vCPU 2 vCPU For comparison: keep-alive at the same level L0 1,374 conn/s 2,038 conn/s 3,802 / 7,863 req/s L6 293 conn/s 741 conn/s 487 / 1,037 req/s Latency, direct path with TLS, GET (p50 / p99 in ms) Measured at 25% · 50% ·…
NullWard documentation