2026-07-29
Bound vLLM request work before it reaches the engine
Define route, byte, fanout, transformation, and tenant work limits that reject oversized vLLM requests before backend allocation.
This article was source-reviewed on 2026-07-29. No local vLLM server, GPU, exploit, benchmark, or production incident was run for this publication. Advisory reproductions are evidence reported by their authors, not results from this article.
One request can represent many work units
A request-rate limit counts arrivals. It does not say how much work one accepted request represents. A completion request can combine an outer prompt list, choices per item, input and output token budgets, structured-output processing, and tenant in-flight work. Each is an independent admission dimension.
The boundary to aim for is rejection before fanout, full materialization, compilation, decoding, or engine scheduling. The n advisory describes synchronous fanout and request-object allocation before scheduling when n was unbounded, and was fixed in 0.19.0. GHSA-3mwp-wvh9-7528 A separate advisory shows why n=1 is not enough: a completion request can still contain an outer prompt list, with each item becoming an engine input, generator, and response slot. GHSA-87x5-vmc3-756j
Built-in vLLM limits are necessary defense in depth. They do not remove the need for an allowlist-first gateway that validates the request shape and enabled feature surface before expensive backend work.
Inventory routes and features before setting limits
Start with a deny-by-default inventory of path, HTTP method, and content type. For every entry, name the owning team, callers, authentication and authorization rule, request-shape policy, and reason it crosses the public boundary. Authentication identifies a caller; it does not bound the work in one authenticated request.
The vLLM 0.26.0 security guide says its API-key middleware does not protect every endpoint and documents inference, operational, utility, development, and profiler surfaces that may be unauthenticated. It recommends a reverse proxy that allowlists needed endpoints and adds authentication, rate limiting, validation, and logging. vLLM 0.26.0 security guide Make an explicit allow or deny decision for inference, operational, utility, development, profiler, speech, derender, prompt-embedding, and structured-output routes and features.
Keep opt-in or disabled-by-default features outside the public boundary unless a documented need, owner, threat review, and admission policy justify them. The prompt-embedding advisory reports that concurrent parts could bypass a process-global sparse invariant guard in that optional feature. Its visible affected and patched ranges are internally inconsistent around 0.26.0, so verify the advisory and exact release rather than treating 0.26.0 as confirmed safe. GHSA-pr7f-p5mw-fc87
Bound five amplification boundaries
Route and feature exposure
Reject unneeded routes and optional capabilities before body processing wherever the gateway can do so. An endpoint with no public use case needs no public parser, validator, or backend path. Prompt embeddings are one example of why feature exposure deserves its own decision rather than being assumed safe because the core inference path is allowed. GHSA-pr7f-p5mw-fc87
Raw bytes and materialization
Apply a streaming raw-body cap before buffering a whole body. For uploads, bound both encoded bytes and the decoded representation. A compressed-size check cannot cap decoded PCM, as the speech advisory explains. GHSA-6pr9-rp53-2pmc Another speech advisory found that the route read a full upload before checking its documented compressed-size limit, fixed in 0.24.0 or later. A limit after materialization does not bound the earlier allocation. GHSA-v82g-2437-67m2
Treat remote media URLs differently from uploaded bytes. Reject them by default. If a route deliberately enables them, require a domain allowlist, revalidate that policy after every redirect, and enforce fetch-byte, timeout, and decoded-representation limits before backend work. vLLM's --allowed-media-domains is a useful defense-in-depth domain control, not a complete gateway budget: it does not replace redirect revalidation or the gateway's byte, timeout, and decoded-representation limits.
Collection fanout
Bound every caller-controlled collection independently: outer prompts, messages, files, parts, choices, and nested generated-output IDs. Reject best_of unless the route deliberately supports and budgets it. If it is supported, cap it independently and use the effective candidate count, at least max(n, best_of), in response-slot and aggregate-work calculations. Add an aggregate item or response-slot budget as well. The outer prompt-list advisory is the direct counterexample to treating n as the only fanout control. Its visible affected and patched metadata are inconsistent, so verify the advisory and exact release rather than treating it as fixed-version guidance. GHSA-87x5-vmc3-756j
Transformation and compile amplification
Byte length is not a complete complexity limit. Gate structured-output backends, and bound or disable attacker-controlled compilation features. The regex advisory reports that structured_outputs.regex could compile attacker-controlled regex without the timeout used by sibling backends when the optional lm-format-enforcer backend was selected. GHSA-48jh-3gj7-fg8v Similarly, caller-supplied generated-output token IDs are ordinary untrusted input, not trusted model output; the derender advisory describes nested IDs accepted without output bounds before decoding and response construction. Its visible metadata is inconsistent around 0.26.0, so do not infer a fixed safe release from it. GHSA-8737-qx52-hjff
Aggregate work and tenant concurrency
Calculate a conservative request budget from all admitted dimensions, then reserve it against a per-tenant in-flight budget before forwarding. One policy can define reserved work as the number of outer items multiplied by the effective candidate count per item, at least max(n, best_of) when the route deliberately supports best_of, multiplied by the sum of the maximum input and output tokens per item, with structured-output and upload-decode surcharges added afterward. A route that does not deliberately support and budget best_of rejects it.
This is a policy-owned approximation, not a universal cost formula. It intentionally overestimates prompt, output, response-slot, structured-output, and upload work so the gateway has a conservative reservation unit. Derive each term from the model, SLO, memory budget, workload, and tenant policy. Cancellation, a client timeout, or a gateway failure must not free admitted work while vLLM may still process it. Release a reservation only after backend-confirmed completion or termination. Otherwise retain the charge through a durable lease, renew it while work is known active, and reconcile conservatively after an uncertain failure.
Gateway admission is separate from benchmark-client concurrency and engine scheduler capacity. --max-num-seqs is not a public request-work contract. See Measure vLLM 0.26.0 tuning before production rollout for that control-boundary distinction.
Write the admission contract as a policy matrix
The following matrix is a contract shape, not a set of defaults. The vLLM guide documents VLLM_MAX_N_SEQUENCES with a default of 16384 and gives 64 or 128 as public-facing examples. That product default and those examples are not deployment recommendations. vLLM 0.26.0 security guide
| Field | Required policy content |
|---|---|
| Route and method | Exact allowlisted path and HTTP method |
| Content type | Exact accepted media types |
| Raw body | <MAX_STREAMED_REQUEST_BYTES> before buffering |
| Outer items | <MAX_OUTER_ITEMS> prompts, messages, files, or parts |
| Choices | <MAX_CHOICES_PER_ITEM> for n or an equivalent output field; reject best_of unless deliberately supported and independently capped |
| Candidate count | If best_of is supported, use at least max(n, best_of) for response slots and aggregate work |
| Input tokens | <MAX_INPUT_TOKENS_PER_ITEM> and <MAX_AGGREGATE_INPUT_TOKENS> |
| Output tokens | <MAX_OUTPUT_TOKENS_PER_ITEM> and <MAX_AGGREGATE_OUTPUT_TOKENS> |
| Aggregate work | <MAX_REQUEST_WORK> using a deployment-defined conservative expression |
| Structured output | <ALLOWED_STRUCTURED_OUTPUT_BACKENDS_AND_CONSTRAINTS> or disabled |
| Upload expansion | <MAX_ENCODED_UPLOAD_BYTES> and <MAX_DECODED_REPRESENTATION_BYTES> |
| Optional features | Explicit allow or deny for embeddings, derender, development, profiler, and operational routes |
| Tenant admission | <MAX_TENANT_RESERVED_IN_FLIGHT_WORK> and <TENANT_REJECTION_POLICY_ID> |
Choose values from the deployed model limits, measured SLO behavior, memory headroom, observed request shapes, tenant contract, and parser behavior. Do not copy an engine default into an internet-facing gateway policy.
Reject in a fixed order
Apply checks in this order:
- Route, method, and content type.
- Streaming raw-byte cap.
- Bounded JSON or multipart parsing.
- Structural dimension caps.
- Aggregate token and response-slot budget.
- Feature-specific transformation or compile policy.
- Atomic tenant work reservation.
- Forward to vLLM.
Earlier checks prevent later amplification. A check after full materialization cannot protect the allocation that already happened. A check on n does not cap an outer prompt list. A raw-byte cap does not bound decoded audio, regex compilation cost, tokenization, or response construction. Keep the implementation contract-neutral: proxy configuration alone cannot establish these guarantees without knowing its parser and buffering behavior.
Prove rejection before the backend
The following is an implementation-neutral test procedure, not commands and not an execution claim. Define a disposable local gateway whose only backend is a stub. The stub counter starts at zero and increments by exactly one for each forwarded request; it makes no vLLM, GPU, or shared-service call. Keep every negative payload small and exceed each local threshold by one unit only.
- For each negative case, snapshot the stub counter before sending the request. Assert a 4xx response, the expected policy ID, and an unchanged counter snapshot afterward. Cover a disallowed route using
DELETEon a small disallowed path, a raw body over its local limit by one byte, outer items over their local limit by one item, aggregate work over its local limit by one unit, and disabled structured output in an otherwise small request. Use the matching route-denied, raw-body-limit, outer-item-limit, aggregate-work-limit, and structured-output-denied policy IDs. - Snapshot the counter, then send one request within every local limit. Assert a 2xx response and exactly one counter increment. This below-limit control shows that the stub is reachable.
- Configure several small same-tenant requests so that their combined reserved work exceeds the shared tenant budget, while each request alone is admissible. Make the stub block each admitted request after reservation until every admission decision and peak-reservation assertion is complete. Hold the requests at a barrier and release them concurrently. For every admitted request, assert that a reservation exists and that total admitted reserved work never exceeds the shared budget. Assert that excess requests receive the expected 4xx tenant-admission policy rejection and do not reach the stub. Then release the admitted stub requests, wait for backend-confirmed termination, and assert that all reservations reconcile to zero.
A rejection test without the backend-counter assertion can pass when the route is broken for an unrelated reason. A control request prevents that false pass. Do not substitute advisory proof-of-concept payloads for these bounded validator tests.
Observe and roll out the contract
Record rejections by policy identifier, route, tenant class, and bounded request-shape metadata. Do not log prompts, credentials, or uploaded content. Compare gateway admitted-work reservations with backend accepted-request counts and vLLM queue or failure observations. An unexplained backend request after a rejection, a reservation leak, a parser-limit mismatch, or a policy bypass is a rejection criterion.
Report-only evaluation is acceptable only when it cannot forward traffic that an existing safety control should already block. Otherwise, enforce first. Then enforce on a bounded traffic slice with predeclared rollback conditions. This rollout observes the admission boundary, not a performance experiment.
Acceptance rule
Accept the admission contract only when every public route and enabled feature has an owner and explicit policy, all caller-controlled expansion dimensions have deployment-derived limits, tenant work is reserved atomically before forwarding, bounded negative tests are rejected without a backend request, below-limit controls still reach the backend, and logs prove the decision without retaining sensitive request content.