undercurrent.adapters¶
Advanced API
ProbedModel picks an adapter for
you with backend="hf" or backend="vllm". Use an adapter directly when
you drive generation yourself, or implement
EngineAdapter to support another
engine.
Importing undercurrent.adapters never loads torch, transformers or vLLM.
EngineAdapter and MissingDependencyError are also exported from
undercurrent; the concrete adapters are imported from their subpackages.
The adapter interface¶
EngineAdapter
¶
Bases: ABC
Abstract base every engine adapter implements.
Implement it to probe an engine Undercurrent doesn't ship an adapter for.
The contract is engine-agnostic: for each request, the caller calls
register_extraction, then generate, then unregister_extraction.
During generate, the adapter builds an
ActivationRecord for every
token/layer that matches one of the request's extraction points and
passes it to router.route(record). When route returns a signal
with action == "abort" for an inline extraction point, the adapter
stops generation as promptly as its engine allows. route already
enforces each extraction point's
InterventionPolicy (including
block_until_signal timeouts), so adapters don't reimplement any
timeout logic; they should document how widely that wait blocks (one
request in a sequential decode loop, every request in the same scheduler
step for a continuously batched engine).
Requests must stay isolated from each other: registering, generating or unregistering one request must not affect another's state.
load_model(model_name_or_path, **kwargs)
abstractmethod
¶
Load the underlying model/engine. Engine-specific kwargs (dtype, device, ...) pass through.
register_extraction(request_id, extraction_points)
abstractmethod
¶
Wire up capture for request_id's extraction points, before generate() is called for it.
Must not affect any other request's state.
generate(request_id, prompt, generation_kwargs, router)
abstractmethod
¶
Run generation for request_id, routing matching activations through router.
Must honor an inline ProbeSignal(action="abort") by stopping
generation as promptly as the underlying engine allows. Returns the
generated text.
unregister_extraction(request_id)
abstractmethod
¶
Tear down whatever register_extraction set up for request_id.
Must be safe to call after generate() returns, normally or by abort,
and must leave no state behind that could leak into a later request.
MissingDependencyError
¶
Bases: ProbingError, ImportError
Raised when an optional package Undercurrent needs isn't installed (most often vLLM).
The message says what to install. A distinct ImportError subclass, so
you can catch it without also swallowing an unrelated import error inside
torch, transformers or vLLM.
Hugging Face transformers¶
HFEngineAdapter()
¶
Bases: EngineAdapter
The EngineAdapter for Hugging Face transformers models.
Captures activations with forward hooks on the decoder layers and their
attention and MLP submodules during model.generate(), and stops
generation through a StoppingCriteria when an inline probe aborts.
An abort takes effect at the next decode step, and a
block_until_signal wait only ever holds up the one request it is for.
Known limitations, which raise
HFAdapterLimitationError:
- one
generate()at a time (no concurrent calls, batch size 1); - no beam search or other decoding that keeps several sequences alive;
- no
tensor_type="kv"; - architectures whose decoder-layer list isn't found (GPT-2-style
transformer.hand Llama-stylemodel.layersare).
adapter = HFEngineAdapter()
adapter.load_model("openai-community/gpt2")
adapter.register_extraction(request_id, extraction_points)
text = adapter.generate(request_id, prompt, {"max_new_tokens": 32}, router)
adapter.unregister_extraction(request_id)
model
property
¶
The loaded transformers model (None before load_model).
tokenizer
property
¶
The loaded tokenizer (None before load_model).
num_layers
property
¶
Number of hookable decoder layers; valid layers indices are 0..num_layers-1.
last_stop(request_id)
¶
(extraction_point_name, signal) of the inline abort that stopped
the most recent generate() call, if that call was for request_id
and was aborted; otherwise None.
close()
¶
Remove every hook this adapter installed on the model and restore
its original forward. The adapter can't generate afterwards.
load_model(model_name_or_path, **kwargs)
¶
Load a model and its tokenizer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_name_or_path
|
Any
|
a Hugging Face hub id or local path (loaded with
|
required |
**kwargs
|
Any
|
|
{}
|
HFAdapterLimitationError
¶
Bases: ProbingError
Raised for a documented, known limitation of the HF adapter.
For example: concurrent generate() calls, an unsupported
tensor_type, or a model architecture whose decoder layers can't be
found. A distinct type, so callers can tell it from a bug.
vLLM¶
VLLMEngineAdapter is public, but the modules behind it read undocumented vLLM internals and are experimental. vLLM is never installed by default; see Deploy with vLLM.
VLLMEngineAdapter()
¶
Bases: EngineAdapter
The EngineAdapter for vLLM (V1 engine), with concurrent generate() calls.
Requires a supported vLLM version, checked on construction (vLLM itself
is imported by load_model). generate() is synchronous, but
load_model starts vLLM's async engine on a background event-loop
thread, so generate() calls from several threads run concurrently and
share scheduler steps. Call
shutdown when
done.
register_extraction() only stores the extraction points; capture is
set up in the worker when generate() has tokenized the prompt.
adapter = VLLMEngineAdapter()
adapter.load_model("openai-community/gpt2")
adapter.register_extraction(request_id, extraction_points)
text = adapter.generate(request_id, prompt, {"max_tokens": 32}, router)
adapter.unregister_extraction(request_id)
See The vLLM adapter for how capture works and its limitations.
shutdown(timeout=30.0)
¶
Shut down the vLLM engine and the background event-loop thread.
Not part of the EngineAdapter contract. Call it when tearing the
adapter down; it's safe to call more than once, and the adapter can't
be used afterwards.
The engine is shut down explicitly, first: vLLM's AsyncLLM runs its
EngineCore in a subprocess and talks to it over ZeroMQ. Leaving that
to garbage collection keeps the subprocess (and its GPU memory) alive,
and collecting the client later can block forever in ZeroMQ's
Context.term() while its sockets are still open. timeout (seconds)
bounds the engine shutdown.
VLLMAdapterLimitationError
¶
Bases: ProbingError
Raised for a documented, known limitation of the vLLM adapter.
For example: an unsupported executor topology, or a vLLM API the adapter doesn't support. A distinct type, so callers can tell it from a bug.