Skip to content

undercurrent.spec

What to capture during generation, expressed as data: extraction points, position selectors and the YAML spec format. The most common names (ExtractionPoint, load_spec, ProbeSpec, the enums) are also exported from undercurrent.

Loading specs

load_spec(spec)

Load a spec from whatever form you have it in.

  • a ProbeSpec is returned unchanged;
  • an os.PathLike (e.g. pathlib.Path) is read with load_yaml_file;
  • a str is a file path if it is one line and names an existing file or ends in .yaml/.yml/.json (read with load_yaml_file); otherwise it is YAML text (parsed with parse_yaml);
  • a mapping is validated with parse_dict.
from undercurrent import load_spec

spec = load_spec("probes.yaml")
spec = load_spec(
    "extraction_points: [{name: p, layers: 0, tensor_type: residual_stream,"
    " position: 'prompt[-1]', probe_type: norm, probe_kind: single_shot}]"
)

Raises:

Type Description
SpecValidationError

the spec is invalid.

SpecFileNotFoundError

the file doesn't exist.

ProbingTypeError

spec is of an unsupported type (a TypeError).

load_yaml_file(path)

Load, parse, and resolve a spec from a YAML file on disk.

parse_yaml(text)

Parse and resolve a spec from a YAML (or JSON, which is valid YAML) string.

parse_dict(data)

Validate a plain dict (already loaded from YAML/JSON) and resolve it.

Raises:

Type Description
SpecValidationError

any structural or semantic problem.

Spec types

ProbeSpec(version, extraction_points) dataclass

A fully resolved spec: an ordered collection of extraction points.

Iterable (yields ExtractionPoints) and sized. Build one with load_spec.

Attributes:

Name Type Description
version str

the spec format version ("1").

extraction_points tuple[ExtractionPoint, ...]

the extraction points, in spec order.

names property

The extraction point names, in spec order.

get(name)

Look up an extraction point by name.

Raises:

Type Description
KeyError

no extraction point has that name.

ExtractionPoint(name, layers, tensor_type, position, stride, until, probe_type, probe_kind, execution_mode, queue_depth, intervention=None, probe_args=FrozenArgs()) dataclass

What to capture, where, and which probe gets it.

One entry of a spec's extraction_points, fully resolved: everything an adapter needs to decide, for each token, whether to capture an activation and which probe to hand it to. Usually built from YAML with load_spec; build one directly with parse_position for position:

ExtractionPoint(
    name="last_prompt_token", layers=(6,), tensor_type=TensorType.RESIDUAL_STREAM,
    position=parse_position("prompt[-1]"), stride=None, until=None,
    probe_type="norm", probe_kind=ProbeKind.SINGLE_SHOT,
    execution_mode=ExecutionMode.INLINE, queue_depth=None,
)

Attributes:

Name Type Description
name str

unique name within the spec; results are keyed by it.

layers tuple[int, ...]

decoder-layer indices to capture at, counted from 0.

tensor_type TensorType

which tensor to capture at each layer.

position PositionSelector

which token positions to capture.

stride int | None

for a continuous position (generated[*] or a slice), capture only every stride-th position. None captures all.

until UntilKind | None

for a continuous position, when to stop capturing (not enforced by the shipped adapters yet; see UntilKind).

probe_type str

the registered name of the probe to run.

probe_kind ProbeKind

the probe's shape; must match the probe's probe_kind.

execution_mode ExecutionMode

inline (can abort) or async (observe only).

queue_depth int | None

async only: how many activations may wait for the probe. None uses the router's default.

intervention InterventionPolicy | None

inline only: the intervention policy. None uses the router's default.

probe_args Mapping[str, Any]

extra keyword arguments for the probe's constructor, merged over the registered factory's kwargs (these win). Stored as a read-only FrozenArgs.

matches(token_index, is_generated, prompt_len, generated_index=None, num_generated_total=None)

Whether this extraction point should fire for a given token.

Combines the position selector match with stride subsampling (which applies only to continuous selectors: 'generated[*]' or a slice). See PositionSelector.matches for the meaning of each argument.

InterventionPolicy(mode=InterventionMode.REJECT, timeout_ms=None, on_timeout=TimeoutAction.CONTINUE) dataclass

How an inline extraction point's probe may hold up generation.

The all-defaults instance (InterventionPolicy()) is the no-op default: mode=reject, timeout_ms=None, on_timeout=continue.

InterventionPolicy(mode="block_until_signal", timeout_ms=50, on_timeout="abort")

Attributes:

Name Type Description
mode InterventionMode
timeout_ms int | None

how long block_until_signal waits for the probe, in milliseconds. Required (>= 1) for that mode, None otherwise.

on_timeout TimeoutAction

what to do when the wait times out or the probe raises; see TimeoutAction.

Raises:

Type Description
ProbingValueError

mode="block_until_signal" without a timeout_ms >= 1.

FrozenArgs(data=None)

Bases: Mapping[str, Any]

Read-only, picklable mapping used for ExtractionPoint.probe_args.

It copies its input, so mutating the dict it was built from doesn't change the extraction point. Compares equal to any mapping with the same items (including a plain dict).

Positions

parse_position(value)

Parse a raw YAML position value into a PositionSelector.

Accepted forms:

Value Selects
5 absolute index 5 in the full (prompt + generated) sequence
"prompt[0]", "prompt[-1]" the first / last prompt token
"generated[0]", "generated[-1]" the first / last generated token
"generated[*]" every generated token (continuous)
"generated[5:]" every generated token from index 5 on (continuous)
"generated[2:8]" generated tokens 2 to 7 (continuous)
"prompt[-1]+1" any form plus an offset: one token past the last prompt token

Raises:

Type Description
PositionSyntaxError

value matches none of the forms; the message includes the accepted grammar.

PositionSelector(kind, index=None, slice_start=None, slice_stop=None, offset=0, raw='') dataclass

Normalized, adapter-facing representation of a position selector.

This is the only representation adapters should consume. Build it once with parse_position and reuse it for every token processed during inference.

Attributes:

Name Type Description
kind PositionKind

which grammar form this selector was parsed from.

index int | None

the raw (possibly negative) index for ABSOLUTE / PROMPT_INDEX / GENERATED_INDEX selectors. Unused otherwise.

slice_start int | None

inclusive start of a GENERATED_SLICE range (always >= 0).

slice_stop int | None

exclusive stop of a GENERATED_SLICE range, or None for an open-ended slice ("generated[5:]").

offset int

an additive offset applied after resolving the base index, e.g. the "+1" in "prompt[-1]+1" ("read position b at token b+1").

raw int | str

the original YAML value this was parsed from, kept for error messages and lossless round-trip serialization.

is_continuous property

True for selectors that can match more than one token position.

resolve_absolute(prompt_len, num_generated_total=None)

Resolve a single-point selector to an absolute sequence index.

Returns None for continuous selectors (wildcard/slice), which have no single absolute index, or when a negative generated[i] selector is asked to resolve before the total generated length is known (streaming decode without a final count yet).

Parameters:

Name Type Description Default
prompt_len int

number of tokens in the prompt.

required
num_generated_total int | None

total number of generated tokens so far (or at final length), required only to resolve negative generated[i] indices such as generated[-1].

None

matches(token_index, is_generated, prompt_len, generated_index=None, num_generated_total=None)

Whether this selector matches a given token.

Parameters:

Name Type Description Default
token_index int

absolute 0-based index of the token in the full (prompt + generated) sequence.

required
is_generated bool

whether this token was generated (vs. part of the prompt).

required
prompt_len int

number of tokens in the prompt.

required
generated_index int | None

0-based index of this token within the generated portion (required to match GENERATED_SLICE / GENERATED_WILDCARD selectors; derivable as token_index - prompt_len when is_generated).

None
num_generated_total int | None

total generated length so far/final, only needed to resolve selectors with negative generated indices (e.g. generated[-1]) during streaming decode. If this is not supplied and it's needed, the selector will not match until it is supplied (so re-check after generation ends for those selectors, e.g. via resolve_absolute()).

None

to_raw()

Reconstruct the YAML scalar this selector would parse from.

Used by the serializer for spec -> internal repr -> spec round trips.

PositionKind

Bases: str, Enum

The form of a position selector, as parsed into a PositionSelector.

ABSOLUTE = 'absolute' class-attribute instance-attribute

A bare integer, e.g. 5.

PROMPT_INDEX = 'prompt_index' class-attribute instance-attribute

prompt[i].

GENERATED_INDEX = 'generated_index' class-attribute instance-attribute

generated[i].

GENERATED_WILDCARD = 'generated_wildcard' class-attribute instance-attribute

generated[*] (continuous).

GENERATED_SLICE = 'generated_slice' class-attribute instance-attribute

generated[start:] or generated[start:stop] (continuous).

Enums

TensorType

Bases: str, Enum

Which tensor an extraction point captures at each of its layers (tensor_type in a spec).

RESIDUAL_STREAM = 'residual_stream' class-attribute instance-attribute

The decoder layer's output hidden state (the residual stream after the layer).

ATTN_OUT = 'attn_out' class-attribute instance-attribute

The attention block's output.

MLP_OUT = 'mlp_out' class-attribute instance-attribute

The MLP block's output.

KV = 'kv' class-attribute instance-attribute

The KV cache. Part of the spec vocabulary, but no shipped adapter captures it.

FINAL_NORM = 'final_norm' class-attribute instance-attribute

The model-level final norm's output (e.g. Llama's trailing RMSNorm).

Distinct from RESIDUAL_STREAM at the last layer, which is the pre-norm residual. A probe trained on transformers' outputs.hidden_states[-1] / outputs.last_hidden_state was trained on this post-norm tensor.

ProbeKind

Bases: str, Enum

A probe's shape: one activation in, one verdict out (SINGLE_SHOT), or state carried across many activations (TRAJECTORY).

SINGLE_SHOT = 'single_shot' class-attribute instance-attribute

One activation in, one verdict out. Runs inline.

TRAJECTORY = 'trajectory' class-attribute instance-attribute

State carried across many activations. Runs inline or async.

ExecutionMode

Bases: str, Enum

INLINE probes run on the generation path and can abort it; ASYNC probes run on a worker pool and only observe.

INLINE = 'inline' class-attribute instance-attribute

Run on the generation path; signals can abort generation.

ASYNC = 'async' class-attribute instance-attribute

Run on a bounded worker pool; observe only. Needs a trajectory probe.

InterventionMode

Bases: str, Enum

How an inline extraction point's ProbeSignal relates to the engine's decode step.

REJECT = 'reject' class-attribute instance-attribute

The default. The router calls on_activation and returns whatever it produces, with no timeout or fallback. The only mode valid for execution_mode=async.

BLOCK_UNTIL_SIGNAL = 'block_until_signal' class-attribute instance-attribute

The router waits up to timeout_ms for on_activation and applies on_timeout if it doesn't return in time (or raises). Requires timeout_ms, so a probe can never hang generation.

TimeoutAction

Bases: str, Enum

What InterventionPolicy falls back to when timeout_ms elapses (or the probe raises) during block_until_signal dispatch.

CONTINUE = 'continue' class-attribute instance-attribute

Carry on generating as if the probe had returned no signal.

ABORT = 'abort' class-attribute instance-attribute

Stop generation.

UntilKind

Bases: str, Enum

When a continuous extraction point stops capturing (until in a spec).

Validated and round-tripped, but not enforced by the shipped adapters yet: a continuous position runs until generation ends. Use a bounded slice such as generated[0:64] for a fixed window.

GENERATION_END = 'generation_end' class-attribute instance-attribute

Capture until generation ends.

FIXED_COUNT = 'fixed_count' class-attribute instance-attribute

Capture a fixed number of positions.

STOP_TOKEN = 'stop_token' class-attribute instance-attribute

Capture until a stop token is generated.

Activation records

ActivationRecord(request_id, extraction_point_name, layer, token_pos, tensor_type, tensor, is_generated, timestamp=time.time()) dataclass

One captured activation, tied to a single extraction point and token.

Attributes:

Name Type Description
request_id str

identifier of the inference request this activation was captured during (opaque; adapters define its format).

extraction_point_name str

the name of the ExtractionPoint that produced this record.

layer int

the specific layer this activation was captured from (a single int, even if the extraction point specified multiple layers -- one record is emitted per matched layer).

token_pos int

absolute 0-based index of the token this activation corresponds to, in the full (prompt + generated) sequence.

tensor_type str

the tensor kind captured, e.g. "residual_stream" (a TensorType value, stored as a plain str).

tensor Any

the captured activation itself. Array-like (with the HF and vLLM adapters, a CPU torch.Tensor); its shape depends on the tensor type.

is_generated bool

whether token_pos falls in the generated portion of the sequence (False for prompt tokens).

timestamp float

unix timestamp (seconds) of when this activation was captured. Defaults to the time of construction.

metadata()

All fields except tensor, as a plain dict.

Useful for logging, indexing, or any other place that wants to handle activation metadata without needing to know how to serialize the tensor itself (numpy/torch tensors generally aren't directly JSON-serializable).

Serialization and JSON Schema

to_yaml(spec)

Serialize a resolved ProbeSpec to a YAML string.

probe_spec_to_dict(spec)

Convert a resolved ProbeSpec back to a plain (spec-shaped) dict.

extraction_point_to_dict(point)

Convert one resolved extraction point back to a plain (spec-shaped) dict.

json_schema()

Return the JSON Schema (Draft 2020-12) for probe-spec YAML files.

The result is a fresh dict each call, with a deterministic key order, so json.dumps(json_schema(), indent=2) is byte-stable.

Errors

ProbingSpecError

Bases: ProbingError

Base class for all errors raised by undercurrent.spec.

SpecValidationError(message, *, issues=None)

Bases: ProbingSpecError

Raised when a spec fails validation.

Covers both structural problems (wrong types, missing fields) and semantic problems (invalid field combinations, e.g. execution_mode=async on a probe_kind=single_shot extraction point). This is a hard failure, never a warning: an invalid spec must never silently produce a resolved ExtractionPoint.

issues holds one self-contained message per problem (with the file and line when known, e.g. probes.yaml:14: extraction_points[2] 'drift'.position: ...); str(exc) combines them.

PositionSyntaxError

Bases: ProbingSpecError

Raised when a position string does not match any supported form.