undercurrent.spec¶
What to capture during generation, expressed as data: extraction points,
position selectors and the YAML spec format. The most common names
(ExtractionPoint, load_spec, ProbeSpec, the enums) are also exported
from undercurrent.
Loading specs¶
load_spec(spec)
¶
Load a spec from whatever form you have it in.
- a
ProbeSpecis returned unchanged; - an
os.PathLike(e.g.pathlib.Path) is read withload_yaml_file; - a
stris a file path if it is one line and names an existing file or ends in.yaml/.yml/.json(read withload_yaml_file); otherwise it is YAML text (parsed withparse_yaml); - a mapping is validated with
parse_dict.
from undercurrent import load_spec
spec = load_spec("probes.yaml")
spec = load_spec(
"extraction_points: [{name: p, layers: 0, tensor_type: residual_stream,"
" position: 'prompt[-1]', probe_type: norm, probe_kind: single_shot}]"
)
Raises:
| Type | Description |
|---|---|
SpecValidationError
|
the spec is invalid. |
SpecFileNotFoundError
|
the file doesn't exist. |
ProbingTypeError
|
|
load_yaml_file(path)
¶
Load, parse, and resolve a spec from a YAML file on disk.
parse_yaml(text)
¶
Parse and resolve a spec from a YAML (or JSON, which is valid YAML) string.
parse_dict(data)
¶
Validate a plain dict (already loaded from YAML/JSON) and resolve it.
Raises:
| Type | Description |
|---|---|
SpecValidationError
|
any structural or semantic problem. |
Spec types¶
ProbeSpec(version, extraction_points)
dataclass
¶
A fully resolved spec: an ordered collection of extraction points.
Iterable (yields ExtractionPoints)
and sized. Build one with load_spec.
Attributes:
| Name | Type | Description |
|---|---|---|
version |
str
|
the spec format version ( |
extraction_points |
tuple[ExtractionPoint, ...]
|
the extraction points, in spec order. |
ExtractionPoint(name, layers, tensor_type, position, stride, until, probe_type, probe_kind, execution_mode, queue_depth, intervention=None, probe_args=FrozenArgs())
dataclass
¶
What to capture, where, and which probe gets it.
One entry of a spec's extraction_points, fully resolved: everything
an adapter needs to decide, for each token, whether to capture an
activation and which probe to hand it to. Usually built from YAML with
load_spec; build one directly with
parse_position for position:
ExtractionPoint(
name="last_prompt_token", layers=(6,), tensor_type=TensorType.RESIDUAL_STREAM,
position=parse_position("prompt[-1]"), stride=None, until=None,
probe_type="norm", probe_kind=ProbeKind.SINGLE_SHOT,
execution_mode=ExecutionMode.INLINE, queue_depth=None,
)
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
unique name within the spec; results are keyed by it. |
layers |
tuple[int, ...]
|
decoder-layer indices to capture at, counted from 0. |
tensor_type |
TensorType
|
which tensor to capture at each layer. |
position |
PositionSelector
|
which token positions to capture. |
stride |
int | None
|
for a continuous position ( |
until |
UntilKind | None
|
for a continuous position, when to stop capturing (not
enforced by the shipped adapters yet; see
|
probe_type |
str
|
the registered name of the probe to run. |
probe_kind |
ProbeKind
|
the probe's shape; must match the probe's |
execution_mode |
ExecutionMode
|
inline (can abort) or async (observe only). |
queue_depth |
int | None
|
async only: how many activations may wait for the probe. None uses the router's default. |
intervention |
InterventionPolicy | None
|
inline only: the intervention policy. None uses the router's default. |
probe_args |
Mapping[str, Any]
|
extra keyword arguments for the probe's constructor,
merged over the registered factory's kwargs (these win). Stored
as a read-only |
matches(token_index, is_generated, prompt_len, generated_index=None, num_generated_total=None)
¶
Whether this extraction point should fire for a given token.
Combines the position selector match with stride subsampling
(which applies only to continuous selectors: 'generated[*]' or a
slice). See
PositionSelector.matches
for the meaning of each argument.
InterventionPolicy(mode=InterventionMode.REJECT, timeout_ms=None, on_timeout=TimeoutAction.CONTINUE)
dataclass
¶
How an inline extraction point's probe may hold up generation.
The all-defaults instance (InterventionPolicy()) is the no-op
default: mode=reject, timeout_ms=None, on_timeout=continue.
Attributes:
| Name | Type | Description |
|---|---|---|
mode |
InterventionMode
|
see |
timeout_ms |
int | None
|
how long |
on_timeout |
TimeoutAction
|
what to do when the wait times out or the probe raises;
see |
Raises:
| Type | Description |
|---|---|
ProbingValueError
|
|
FrozenArgs(data=None)
¶
Bases: Mapping[str, Any]
Read-only, picklable mapping used for ExtractionPoint.probe_args.
It copies its input, so mutating the dict it was built from doesn't change the extraction point. Compares equal to any mapping with the same items (including a plain dict).
Positions¶
parse_position(value)
¶
Parse a raw YAML position value into a PositionSelector.
Accepted forms:
| Value | Selects |
|---|---|
5 |
absolute index 5 in the full (prompt + generated) sequence |
"prompt[0]", "prompt[-1]" |
the first / last prompt token |
"generated[0]", "generated[-1]" |
the first / last generated token |
"generated[*]" |
every generated token (continuous) |
"generated[5:]" |
every generated token from index 5 on (continuous) |
"generated[2:8]" |
generated tokens 2 to 7 (continuous) |
"prompt[-1]+1" |
any form plus an offset: one token past the last prompt token |
Raises:
| Type | Description |
|---|---|
PositionSyntaxError
|
|
PositionSelector(kind, index=None, slice_start=None, slice_stop=None, offset=0, raw='')
dataclass
¶
Normalized, adapter-facing representation of a position selector.
This is the only representation adapters should consume. Build it once
with parse_position and reuse it
for every token processed during inference.
Attributes:
| Name | Type | Description |
|---|---|---|
kind |
PositionKind
|
which grammar form this selector was parsed from. |
index |
int | None
|
the raw (possibly negative) index for ABSOLUTE / PROMPT_INDEX / GENERATED_INDEX selectors. Unused otherwise. |
slice_start |
int | None
|
inclusive start of a GENERATED_SLICE range (always >= 0). |
slice_stop |
int | None
|
exclusive stop of a GENERATED_SLICE range, or None for an open-ended slice ("generated[5:]"). |
offset |
int
|
an additive offset applied after resolving the base index, e.g. the "+1" in "prompt[-1]+1" ("read position b at token b+1"). |
raw |
int | str
|
the original YAML value this was parsed from, kept for error messages and lossless round-trip serialization. |
is_continuous
property
¶
True for selectors that can match more than one token position.
resolve_absolute(prompt_len, num_generated_total=None)
¶
Resolve a single-point selector to an absolute sequence index.
Returns None for continuous selectors (wildcard/slice), which have
no single absolute index, or when a negative generated[i]
selector is asked to resolve before the total generated length is
known (streaming decode without a final count yet).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prompt_len
|
int
|
number of tokens in the prompt. |
required |
num_generated_total
|
int | None
|
total number of generated tokens so far
(or at final length), required only to resolve negative
|
None
|
matches(token_index, is_generated, prompt_len, generated_index=None, num_generated_total=None)
¶
Whether this selector matches a given token.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token_index
|
int
|
absolute 0-based index of the token in the full (prompt + generated) sequence. |
required |
is_generated
|
bool
|
whether this token was generated (vs. part of the prompt). |
required |
prompt_len
|
int
|
number of tokens in the prompt. |
required |
generated_index
|
int | None
|
0-based index of this token within the
generated portion (required to match GENERATED_SLICE /
GENERATED_WILDCARD selectors; derivable as
|
None
|
num_generated_total
|
int | None
|
total generated length so far/final, only
needed to resolve selectors with negative generated indices
(e.g. |
None
|
to_raw()
¶
Reconstruct the YAML scalar this selector would parse from.
Used by the serializer for spec -> internal repr -> spec round trips.
PositionKind
¶
Bases: str, Enum
The form of a position selector, as parsed into a PositionSelector.
ABSOLUTE = 'absolute'
class-attribute
instance-attribute
¶
A bare integer, e.g. 5.
PROMPT_INDEX = 'prompt_index'
class-attribute
instance-attribute
¶
prompt[i].
GENERATED_INDEX = 'generated_index'
class-attribute
instance-attribute
¶
generated[i].
GENERATED_WILDCARD = 'generated_wildcard'
class-attribute
instance-attribute
¶
generated[*] (continuous).
GENERATED_SLICE = 'generated_slice'
class-attribute
instance-attribute
¶
generated[start:] or generated[start:stop] (continuous).
Enums¶
TensorType
¶
Bases: str, Enum
Which tensor an extraction point captures at each of its layers (tensor_type in a spec).
RESIDUAL_STREAM = 'residual_stream'
class-attribute
instance-attribute
¶
The decoder layer's output hidden state (the residual stream after the layer).
ATTN_OUT = 'attn_out'
class-attribute
instance-attribute
¶
The attention block's output.
MLP_OUT = 'mlp_out'
class-attribute
instance-attribute
¶
The MLP block's output.
KV = 'kv'
class-attribute
instance-attribute
¶
The KV cache. Part of the spec vocabulary, but no shipped adapter captures it.
FINAL_NORM = 'final_norm'
class-attribute
instance-attribute
¶
The model-level final norm's output (e.g. Llama's trailing RMSNorm).
Distinct from RESIDUAL_STREAM at the last layer, which is the
pre-norm residual. A probe trained on transformers'
outputs.hidden_states[-1] / outputs.last_hidden_state was
trained on this post-norm tensor.
ProbeKind
¶
Bases: str, Enum
A probe's shape: one activation in, one verdict out (SINGLE_SHOT), or
state carried across many activations (TRAJECTORY).
ExecutionMode
¶
Bases: str, Enum
INLINE probes run on the generation path and can abort it; ASYNC
probes run on a worker pool and only observe.
InterventionMode
¶
Bases: str, Enum
How an inline extraction point's ProbeSignal relates to the
engine's decode step.
REJECT = 'reject'
class-attribute
instance-attribute
¶
The default. The router calls on_activation and returns whatever
it produces, with no timeout or fallback. The only mode valid for
execution_mode=async.
BLOCK_UNTIL_SIGNAL = 'block_until_signal'
class-attribute
instance-attribute
¶
The router waits up to timeout_ms for on_activation and
applies on_timeout if it doesn't return in time (or raises).
Requires timeout_ms, so a probe can never hang generation.
TimeoutAction
¶
Bases: str, Enum
What InterventionPolicy falls back to when timeout_ms elapses
(or the probe raises) during block_until_signal dispatch.
UntilKind
¶
Bases: str, Enum
When a continuous extraction point stops capturing (until in a spec).
Validated and round-tripped, but not enforced by the shipped adapters
yet: a continuous position runs until generation ends. Use a bounded
slice such as generated[0:64] for a fixed window.
GENERATION_END = 'generation_end'
class-attribute
instance-attribute
¶
Capture until generation ends.
FIXED_COUNT = 'fixed_count'
class-attribute
instance-attribute
¶
Capture a fixed number of positions.
STOP_TOKEN = 'stop_token'
class-attribute
instance-attribute
¶
Capture until a stop token is generated.
Activation records¶
ActivationRecord(request_id, extraction_point_name, layer, token_pos, tensor_type, tensor, is_generated, timestamp=time.time())
dataclass
¶
One captured activation, tied to a single extraction point and token.
Attributes:
| Name | Type | Description |
|---|---|---|
request_id |
str
|
identifier of the inference request this activation was captured during (opaque; adapters define its format). |
extraction_point_name |
str
|
the |
layer |
int
|
the specific layer this activation was captured from (a single int, even if the extraction point specified multiple layers -- one record is emitted per matched layer). |
token_pos |
int
|
absolute 0-based index of the token this activation corresponds to, in the full (prompt + generated) sequence. |
tensor_type |
str
|
the tensor kind captured, e.g. |
tensor |
Any
|
the captured activation itself. Array-like (with the HF and
vLLM adapters, a CPU |
is_generated |
bool
|
whether |
timestamp |
float
|
unix timestamp (seconds) of when this activation was captured. Defaults to the time of construction. |
metadata()
¶
All fields except tensor, as a plain dict.
Useful for logging, indexing, or any other place that wants to handle activation metadata without needing to know how to serialize the tensor itself (numpy/torch tensors generally aren't directly JSON-serializable).
Serialization and JSON Schema¶
to_yaml(spec)
¶
Serialize a resolved ProbeSpec to a YAML string.
probe_spec_to_dict(spec)
¶
Convert a resolved ProbeSpec back to a plain (spec-shaped) dict.
extraction_point_to_dict(point)
¶
Convert one resolved extraction point back to a plain (spec-shaped) dict.
json_schema()
¶
Return the JSON Schema (Draft 2020-12) for probe-spec YAML files.
The result is a fresh dict each call, with a deterministic key order,
so json.dumps(json_schema(), indent=2) is byte-stable.
Errors¶
ProbingSpecError
¶
Bases: ProbingError
Base class for all errors raised by undercurrent.spec.
SpecValidationError(message, *, issues=None)
¶
Bases: ProbingSpecError
Raised when a spec fails validation.
Covers both structural problems (wrong types, missing fields) and
semantic problems (invalid field combinations, e.g.
execution_mode=async on a probe_kind=single_shot extraction
point). This is a hard failure, never a warning: an invalid spec must
never silently produce a resolved ExtractionPoint.
issues holds one self-contained message per problem (with the file
and line when known, e.g. probes.yaml:14: extraction_points[2]
'drift'.position: ...); str(exc) combines them.
PositionSyntaxError
¶
Bases: ProbingSpecError
Raised when a position string does not match any supported form.