Skip to content

undercurrent.adapters

Advanced API

ProbedModel picks an adapter for you with backend="hf" or backend="vllm". Use an adapter directly when you drive generation yourself, or implement EngineAdapter to support another engine.

Importing undercurrent.adapters never loads torch, transformers or vLLM. EngineAdapter and MissingDependencyError are also exported from undercurrent; the concrete adapters are imported from their subpackages.

The adapter interface

EngineAdapter

Bases: ABC

Abstract base every engine adapter implements.

Implement it to probe an engine Undercurrent doesn't ship an adapter for. The contract is engine-agnostic: for each request, the caller calls register_extraction, then generate, then unregister_extraction.

During generate, the adapter builds an ActivationRecord for every token/layer that matches one of the request's extraction points and passes it to router.route(record). When route returns a signal with action == "abort" for an inline extraction point, the adapter stops generation as promptly as its engine allows. route already enforces each extraction point's InterventionPolicy (including block_until_signal timeouts), so adapters don't reimplement any timeout logic; they should document how widely that wait blocks (one request in a sequential decode loop, every request in the same scheduler step for a continuously batched engine).

Requests must stay isolated from each other: registering, generating or unregistering one request must not affect another's state.

load_model(model_name_or_path, **kwargs) abstractmethod

Load the underlying model/engine. Engine-specific kwargs (dtype, device, ...) pass through.

register_extraction(request_id, extraction_points) abstractmethod

Wire up capture for request_id's extraction points, before generate() is called for it.

Must not affect any other request's state.

generate(request_id, prompt, generation_kwargs, router) abstractmethod

Run generation for request_id, routing matching activations through router.

Must honor an inline ProbeSignal(action="abort") by stopping generation as promptly as the underlying engine allows. Returns the generated text.

unregister_extraction(request_id) abstractmethod

Tear down whatever register_extraction set up for request_id.

Must be safe to call after generate() returns, normally or by abort, and must leave no state behind that could leak into a later request.

MissingDependencyError

Bases: ProbingError, ImportError

Raised when an optional package Undercurrent needs isn't installed (most often vLLM).

The message says what to install. A distinct ImportError subclass, so you can catch it without also swallowing an unrelated import error inside torch, transformers or vLLM.

Hugging Face transformers

from undercurrent.adapters.hf import HFEngineAdapter

HFEngineAdapter()

Bases: EngineAdapter

The EngineAdapter for Hugging Face transformers models.

Captures activations with forward hooks on the decoder layers and their attention and MLP submodules during model.generate(), and stops generation through a StoppingCriteria when an inline probe aborts. An abort takes effect at the next decode step, and a block_until_signal wait only ever holds up the one request it is for.

Known limitations, which raise HFAdapterLimitationError:

  • one generate() at a time (no concurrent calls, batch size 1);
  • no beam search or other decoding that keeps several sequences alive;
  • no tensor_type="kv";
  • architectures whose decoder-layer list isn't found (GPT-2-style transformer.h and Llama-style model.layers are).
adapter = HFEngineAdapter()
adapter.load_model("openai-community/gpt2")
adapter.register_extraction(request_id, extraction_points)
text = adapter.generate(request_id, prompt, {"max_new_tokens": 32}, router)
adapter.unregister_extraction(request_id)

model property

The loaded transformers model (None before load_model).

tokenizer property

The loaded tokenizer (None before load_model).

num_layers property

Number of hookable decoder layers; valid layers indices are 0..num_layers-1.

last_stop(request_id)

(extraction_point_name, signal) of the inline abort that stopped the most recent generate() call, if that call was for request_id and was aborted; otherwise None.

close()

Remove every hook this adapter installed on the model and restore its original forward. The adapter can't generate afterwards.

load_model(model_name_or_path, **kwargs)

Load a model and its tokenizer.

Parameters:

Name Type Description Default
model_name_or_path Any

a Hugging Face hub id or local path (loaded with AutoModelForCausalLM / AutoTokenizer), or an already-built model object, which then needs tokenizer=.

required
**kwargs Any

device= (default "cpu") and tokenizer=; the rest is passed to from_pretrained (e.g. torch_dtype).

{}

HFAdapterLimitationError

Bases: ProbingError

Raised for a documented, known limitation of the HF adapter.

For example: concurrent generate() calls, an unsupported tensor_type, or a model architecture whose decoder layers can't be found. A distinct type, so callers can tell it from a bug.

vLLM

from undercurrent.adapters.vllm import VLLMEngineAdapter

VLLMEngineAdapter is public, but the modules behind it read undocumented vLLM internals and are experimental. vLLM is never installed by default; see Deploy with vLLM.

VLLMEngineAdapter()

Bases: EngineAdapter

The EngineAdapter for vLLM (V1 engine), with concurrent generate() calls.

Requires a supported vLLM version, checked on construction (vLLM itself is imported by load_model). generate() is synchronous, but load_model starts vLLM's async engine on a background event-loop thread, so generate() calls from several threads run concurrently and share scheduler steps. Call shutdown when done.

register_extraction() only stores the extraction points; capture is set up in the worker when generate() has tokenized the prompt.

adapter = VLLMEngineAdapter()
adapter.load_model("openai-community/gpt2")
adapter.register_extraction(request_id, extraction_points)
text = adapter.generate(request_id, prompt, {"max_tokens": 32}, router)
adapter.unregister_extraction(request_id)

See The vLLM adapter for how capture works and its limitations.

shutdown(timeout=30.0)

Shut down the vLLM engine and the background event-loop thread.

Not part of the EngineAdapter contract. Call it when tearing the adapter down; it's safe to call more than once, and the adapter can't be used afterwards.

The engine is shut down explicitly, first: vLLM's AsyncLLM runs its EngineCore in a subprocess and talks to it over ZeroMQ. Leaving that to garbage collection keeps the subprocess (and its GPU memory) alive, and collecting the client later can block forever in ZeroMQ's Context.term() while its sockets are still open. timeout (seconds) bounds the engine shutdown.

VLLMAdapterLimitationError

Bases: ProbingError

Raised for a documented, known limitation of the vLLM adapter.

For example: an unsupported executor topology, or a vLLM API the adapter doesn't support. A distinct type, so callers can tell it from a bug.