Deploy with vLLM¶
This page covers running Undercurrent on vLLM:
installing it next to the vLLM you already run, generating with
ProbedModel(..., backend="vllm"), returning probe verdicts over HTTP, and
which vLLM versions are supported. If you have your own request loop or
server and want to drive the engine adapter and Router yourself, read
Embed in your serving stack after this page.
Requires a GPU
Every vLLM example on this page needs a CUDA GPU and a working vLLM install. None of them run in the docs CI. To try Undercurrent without a GPU, use the default Hugging Face backend from the Quickstart.
Known issues (v0.1)
The vLLM adapter hasn't been verified on a GPU yet. Fixed in 0.1.0: on
Llama-family models (and other fused-residual architectures),
residual_stream is captured as hidden_states + residual, the same
residual stream as the HF backend; earlier builds captured the layer's
MLP output. GPU verification of the fix is pending. See
Known issues for details.
Install into your vLLM environment¶
vLLM ships builds matched to a specific torch and CUDA version. Undercurrent doesn't try to pick those for you. The recommended production setup is to install Undercurrent into an environment or image that already has a working, GPU-matched vLLM:
For your own container image, start from the vLLM image you already use and add Undercurrent on top:
# Pick a vllm/vllm-openai tag whose vLLM is in Undercurrent's supported range
# (see docs/compatibility.md).
FROM vllm/vllm-openai:<tag>
RUN pip install undercurrent
# The base image's entrypoint starts vLLM's own OpenAI-compatible server, which
# doesn't drive probes (see "Serving over HTTP" below). Run your own server.
COPY server.py /app/server.py
ENTRYPOINT ["python3", "/app/server.py"]
The snippet is a starting point for your own image; adapt the tag, the entrypoint and how you install your server's code.
Undercurrent never installs or upgrades vLLM or torch behind your back.
vLLM isn't a dependency of pip install undercurrent at all, and the torch
requirement is a loose floor (torch>=2.1, no upper bound), so it never
fights the exact torch version your vLLM pins. Installing Undercurrent into a
working vLLM environment leaves vLLM and torch as they were. The project's
GPU CI installs vLLM first, then Undercurrent, and fails if either version
changed.
The convenience extra. If you're starting from an empty environment, for example on a development GPU box, you can let pip install a vLLM from the tested range:
This resolves vllm>=0.28,<0.29 together with whatever torch that vLLM
wheel requires. Use it when you don't already have a vLLM to match. For
production images, prefer the vLLM install (or base image) you already trust,
and add pip install undercurrent to it.
The runtime version check. Because pip never sees the supported range
when you install into an existing environment, the adapter checks it itself.
Constructing a VLLMEngineAdapter, which ProbedModel(backend="vllm") does
for you, reads the installed vLLM version from package metadata (it doesn't
import vLLM) and compares it with SUPPORTED_VLLM (>=0.28,<0.29). Outside
that range it raises VLLMAdapterLimitationError before any model is
loaded; see Troubleshooting. If vLLM isn't installed at
all, you get MissingDependencyError instead, listing the same two install
routes.
Compatibility lists the declared ranges, the combinations that have actually been tested, and how to override the check.
Generate with ProbedModel¶
ProbedModel is the same front door as in the
Quickstart; pass backend="vllm".
Keyword arguments that ProbedModel doesn't use itself go to vLLM as engine
arguments (gpu_memory_utilization, max_model_len, dtype,
trust_remote_code, ...), together with the adapter's own switches such as
allow_unsupported_executor.
Requires a GPU
This example needs a CUDA GPU and vLLM.
from undercurrent.core import ActivationRecord, probe
from undercurrent.model import ProbedModel
SPEC = """
version: "1"
extraction_points:
- name: last_prompt_token
layers: 6
tensor_type: residual_stream
position: "prompt[-1]"
probe_type: norm_gate
probe_kind: single_shot
execution_mode: inline
"""
@probe("norm_gate", threshold=500.0)
def norm_gate(record: ActivationRecord) -> float:
return float(record.tensor.float().norm())
prompts = ["Hello, my name is", "The capital of France is", "Once upon a time"]
with ProbedModel.from_pretrained(
"openai-community/gpt2",
backend="vllm",
spec=SPEC,
probes={"norm_gate": norm_gate},
gpu_memory_utilization=0.3, # vLLM engine args from here on
max_model_len=1024,
enforce_eager=True, # needed for capture; see "Scaling and limits"
) as model:
outputs = model.generate(prompts, max_new_tokens=32, temperature=0)
for out in outputs:
print(out.request_id, repr(out.text), out.probe_results["last_prompt_token"].verdict)
What's different from the HF backend:
- Batching.
generate()on a list runs up tomax_concurrencyprompts at once (default 64,DEFAULT_VLLM_MAX_CONCURRENCY; set it withProbedModel(..., max_concurrency=...)), and vLLM's scheduler batches them. The outputs come back as a list in prompt order. Each prompt is its own request with its ownrequest_idand its own probe instances, so results never mix between prompts. - Failures in a batch. By default the first failing prompt raises for the
whole call, and prompts that haven't started are cancelled. Pass
return_exceptions=Trueto get the exception in that prompt's slot of the returned list instead. - Generation arguments.
max_new_tokensbecomes vLLM'smax_tokens(default 32 when you don't give one),min_new_tokensbecomesmin_tokens, andtemperature(0 is greedy),top_p,seedandstopmap one to one ontoSamplingParams. Any other keyword argument is passed through as aSamplingParamsfield. - Device. vLLM places the model on its GPU(s) itself. Leave
device=unset; anything other thanNoneor acudadevice raisesProbedModelConfigError. - Tensor types.
kvisn't supported on vLLM.ProbedModelrejects it, and layer indices that are out of range for the model, at construction, before the model loads.
Stream results into your own system: on_result=¶
on_result is called once per finished generation with its
GenerationOutput, as soon as that prompt finishes, not when the whole batch
is done. With the vLLM backend it runs on the thread that finished the
generation, one of the batch's worker threads, so several calls can run at
the same time. Keep the callback thread-safe and quick, for example by
putting the output on a queue that another thread ships to your database or
message bus. If the callback raises, the exception is logged on the
undercurrent logger and the generation's output is still returned.
Requires a GPU
This example needs a CUDA GPU and vLLM.
import queue
results: queue.Queue = queue.Queue()
with ProbedModel.from_pretrained(
"openai-community/gpt2",
backend="vllm",
spec=SPEC,
probes={"norm_gate": norm_gate},
on_result=results.put, # queue.Queue.put is thread-safe
enforce_eager=True,
) as model:
model.generate(prompts, max_new_tokens=32)
while not results.empty():
out = results.get()
print(out.request_id, out.aborted, out.probe_results["last_prompt_token"].verdict)
For durable, out-of-process delivery of async (observe-mode) probe output,
use a log sink instead (ProbedModel(..., log_sink=...)); see
Observation sinks.
What an inline abort looks like¶
When an inline probe returns an ABORT signal, the adapter cancels the
request through vLLM's own abort API. The generation's GenerationOutput
then has:
| Field | Value after an abort |
|---|---|
aborted |
True |
abort_point |
Name of the extraction point that stopped generation, for example "last_prompt_token" |
abort_signal |
The ProbeSignal that stopped it (its confidence and metadata) |
abort_reason |
A readable summary, for example extraction point 'last_prompt_token' aborted generation (confidence 0.97) |
text |
The text generated before vLLM stopped the request |
probe_results |
Every extraction point's ProbeResult, the aborting one included |
Two things differ from the HF backend:
- The abort lands a step late. vLLM schedules a whole step before it
checks for aborted requests, so a request stops at the start of the next
scheduler step. Expect at least one extra token after the one whose
activation triggered the abort, and a few more under heavy concurrent load.
Don't return
textfrom an aborted generation to your users if the abort is a safety gate. abortedcomes from the probe's results. The vLLM adapter doesn't report which extraction point stopped a request, soProbedModelfinds the earliestABORTsignal in the inline extraction points'ProbeResult.signal_history. Function probes (@probe) record their signals there automatically. A class probe must include theABORTsignals it returned in thesignal_historyof theProbeResultfromon_end, as the Quickstart'sMeanNormProbedoes. Otherwise vLLM still stops the request, butabortedisFalse.
for out in outputs:
if out.aborted:
print(f"{out.request_id}: stopped by {out.abort_point}: {out.abort_reason}")
else:
print(f"{out.request_id}: {out.text!r}")
Bring your own adapter or router¶
Everything below ProbedModel stays reachable:
router_kwargs={...}is forwarded toRouter(...):worker_pool_size,default_queue_depth,default_overflow_policy,drain_timeout,default_intervention_policy,circuit_breaker_threshold. See Embed in your serving stack for what to set.log_sink=andmetrics_sink=attach an observation sink and a metrics sink.backend=VLLMEngineAdapter()uses an adapter you built and loaded yourself (passNoneas the model). You keep owning it:close()doesn't shut it down.model.routerandmodel.adapterexpose the live objects.
Serving over HTTP¶
Undercurrent v0.1 doesn't include an HTTP server. A supported, OpenAI-compatible
serve command is on the roadmap.
Until then, put VLLMEngineAdapter (or ProbedModel) behind the web server
you already run.
examples/openai_server/
is the reference to copy. It isn't installed with the package. It has two
parts:
serve_llama_mlp_pipeline.py: a working HTTP server. It builds oneRouterand oneVLLMEngineAdapterat startup, shares them across requests, and handles each HTTP request on its own thread withregister_extraction→generate→unregister_extraction. It collects each request's probe results withrouter.on_request_end. It uses the stdlibThreadingHTTPServer, which isn't hardened for production (no backpressure, request timeouts or TLS); keep itsrun_generate()logic and put it behind uvicorn, gunicorn or your own framework.response_schema.py: a pure formatting layer, with no HTTP or vLLM dependency, that turns a finished generation and its probe verdicts into OpenAI-style response bodies and SSE chunks.
To run the reference server (it needs a probe checkpoint; the README and the script's docstring cover making a placeholder one):
export PROBE_MODEL_PATH=/path/to/checkpoint.pt
python examples/openai_server/serve_llama_mlp_pipeline.py --port 8000
The server listens on 127.0.0.1 by default. In a container, pass
--host 0.0.0.0 so the published port can reach it. The server has no
authentication, so keep it off untrusted networks.
The reference server returns its own simple shape: request_id, prompt,
text, and a verdicts object with each extraction point's
ProbeResult.verdict. To return the OpenAI-style wire format instead, build
the body with response_schema.build_completion_response(outcome, verdicts)
from a GenerationOutcome and one ProbeVerdict per extraction point.
A clean completion carries a probing object with one entry per extraction
point, inline and async alike:
{
"flagged": false,
"finish_reason": "stop",
"text": "hello world",
"tokens_generated": 42,
"max_tokens": 256,
"probing": {
"toxicity_probe": {
"score": 0.12,
"flagged": false,
"execution_mode": "inline",
"evaluated": true
},
"trajectory_probe": {
"score": {"max_score": 0.4, "final_score": 0.2},
"flagged": false,
"execution_mode": "async",
"evaluated": true
}
}
}
scoreis a float for a single-shot probe, or{max_score, final_score}for a trajectory probe (aggregate_trajectory_scores).evaluatedisfalsewhen the extraction point's position was never reached, for example agenerated[*]point on a generation that produced no tokens. Thenscoreisnull.
When an inline probe aborted the generation, the body has a different shape.
The text is suppressed and finish_reason is "content_flagged" instead of
"stop":
{
"flagged": true,
"finish_reason": "content_flagged",
"text": null,
"flagged_by": "toxicity_probe",
"flagged_score": 0.94,
"tokens_generated": 17,
"max_tokens": 256,
"tokens_saved": 239
}
flagged_byandflagged_scorecome from the first inline, evaluated, flagged verdict. An async probe is never reported as the cause, because async probes never stop generation.tokens_savedismax_tokens - tokens_generated.
For streaming, build_flagged_sse_chunk() builds the terminal chunk of a
flagged stream (choices[0].finish_reason == "content_flagged", empty
delta), and build_probing_result_sse_event() builds an extra
probing_result event to send after a clean stream's normal final chunk.
Scaling and limits¶
- One worker per engine. The adapter is validated with a single vLLM
worker. If vLLM starts more than one worker, for tensor or pipeline
parallelism,
load_model()raisesVLLMAdapterLimitationErrorunless you passallow_unsupported_executor=True. With the override, the adapter merges every worker's records and drops tensor-parallel duplicates, but multi-GPU capture isn't validated yet. To scale out today, run several single-GPU engines behind a load balancer. Tensor & pipeline parallelism has the per-topology support status. - Use
enforce_eager=True. With CUDA graphs, replayed steps don't run Python, so the forward hooks can't fire. The adapter doesn't set it for you. block_until_signalstalls the whole batch. Under continuous batching, a blocking inline probe holds the scheduler step that every co-batched request shares, for up totimeout_ms. The adapter logs a one-time warning when it sees such an extraction point. Useblock_until_signalon vLLM only for single-request or offline runs; with concurrent traffic, keep inline probes cheap with the defaultmode: reject, or run themexecution_mode: async. See Intervention policies.- Model coverage. Some captures silently produce nothing or something
other than you'd expect on particular architectures (for example
attn_outon GPT-2-style models, orresidual_streamon Llama-family layers). Check Known limitations for your model before relying on a probe. vllm serve. Undercurrent registers avllm.general_pluginsentry point, but a barevllm serveprocess doesn't drive probing. See Running undervllm serve.
The vLLM adapter explains how capture, position mapping and aborts work.
Supported versions¶
The supported vLLM range is a single minor release, currently
vllm>=0.28,<0.29. The adapter hooks undocumented vLLM internals that can
change in any minor release, so the range stops at the next one.
- Pinned minor. This is the range the GPU test suite targets. Before
every release a maintainer runs it on a GPU machine with
scripts/gpu_check.sh, installed the documented way: vLLM first, then Undercurrent into the same environment. - Moving to a new minor. The range is widened, in the
[vllm]extra,SUPPORTED_VLLMand Compatibility together, only after the GPU suite passes on the new minor. The previous minor stays best-effort for one more Undercurrent release, then it's dropped.
GPU CI status
There is no GPU CI for v0.1: the GPU tests run by hand before each release
(scripts/gpu_check.sh). The vLLM rows on the
Compatibility page say which
combinations have had a GPU run.
Troubleshooting¶
VLLMAdapterLimitationError: undercurrent.adapters.vllm supports vllm>=0.28,<0.29, but vllm==X is installed. ...
Raised when the adapter is constructed. Your vLLM is outside the supported
range. Either use an environment or base image with a vLLM in range, or
install the tested range with pip install "undercurrent[vllm]" in a fresh
environment. To try an untested version anyway, for example to evaluate a
new vLLM release, set UNDERCURRENT_ALLOW_UNSUPPORTED_VLLM=1 (true,
yes and on also work). The error becomes a RuntimeWarning, and the
checks below still fail loudly if vLLM's internals changed.
MissingDependencyError: backend='vllm' needs vLLM, which undercurrent never installs by default. ...
vLLM isn't importable in this environment. Install Undercurrent into your
vLLM environment, or use pip install "undercurrent[vllm]".
VLLMAdapterLimitationError: detected N workers (via collective_rpc's one-result-per-worker fan-out). ...
Your engine runs more than one worker, usually because
tensor_parallel_size or pipeline_parallel_size is above 1. Run one
worker per engine, or pass allow_unsupported_executor=True if you accept
an unvalidated topology (see Scaling and limits).
VLLMAdapterLimitationError: could not determine worker topology (a probe RPC call raised ...)
The adapter's worker-count check failed. This usually means the installed
vLLM's RPC surface differs from what the adapter expects. Check the vLLM
version against Compatibility. If you're certain the
engine has exactly one worker, allow_unsupported_executor=True skips the
check.
VLLMAdapterLimitationError: could not find a collective_rpc() entry point on the loaded engine ... or ... could not locate a tokenizer on the loaded engine ...
The engine object doesn't look like the vLLM the adapter was written against. This happens with an unsupported vLLM version, typically one loaded with the override. Use a supported version.
VLLMAdapterLimitationError: a different Router instance was already bound to this adapter ...
One VLLMEngineAdapter serves exactly one Router for its whole lifetime.
Build one Router at startup and pass the same instance to every
generate() call.
VLLMIntrospectionError: model_runner.input_batch is missing ..., ... is missing expected attribute(s) [...]
Raised on the first forward hook when vLLM's internal batch layout isn't
the V1 model-runner layout the adapter reads. Causes, most likely first:
an unsupported vLLM version (with the version check overridden), or
allow_v2_model_runner=True (the adapter only understands the V1 model
runner; keep the default, which forces V1). The adapter fails here on
purpose rather than attach activations to the wrong tokens.
SeqMapperError: ... has a scheduled row range but was never registered via register_request()
vLLM reported a request id the adapter didn't register. This happens when
vLLM rewrites request ids, which the adapter prevents by setting
VLLM_DISABLE_REQUEST_ID_RANDOMIZATION=1 in load_model(). You see this
error if you passed allow_request_id_randomization=True, or if your vLLM
no longer honours that variable (an unsupported version).
WorkerExtractionError: could not locate a decoder-layer list on the loaded model ...
The model's architecture keeps its decoder layers somewhere the adapter
doesn't look (it checks model.model.layers, model.layers and
model.transformer.h). Use a supported architecture, or add the path in
ProbingWorkerExtension._find_decoder_layers() and open a pull request.
No probe results, or aborted is False although a probe asked to abort
Check that you passed enforce_eager=True, that the extraction point's
tensor_type and position are captured on your model (see
Known limitations), and
that a class probe puts its ABORT signals in signal_history (see
What an inline abort looks like).