Skip to content

undercurrent.core

How a probe plugs in: the probe lifecycle, function probes, the probe registry, and the signal and result types probes produce. The front-door names (probe, register_probe, Probe) and the data types are also exported from undercurrent.

Writing probes

Decorate a plain function with @probe for a stateless single-shot probe, or subclass Probe for stateful and trajectory probes. @register_probe makes a class findable by the probe_type name used in specs.

probe(name=None, /, *, threshold=None, action='abort', registry=None, register=True)

probe(name: ProbeFunction) -> type[FunctionProbe]
probe(name: str | None = None, /, *, threshold: float | None = None, action: str = 'abort', registry: ProbeRegistry | None = None, register: bool = True) -> Callable[[ProbeFunction], type[FunctionProbe]]

Turn fn(record) -> float | bool | ProbeSignal | None into a single_shot Probe.

@probe("norm", threshold=10.0)
def norm(record, *, ord: int = 2) -> float:
    return float(record.tensor.float().norm(p=ord))

What the function returns, per activation:

  • None: nothing to report. No signal, not counted.
  • bool: True flags the activation, False doesn't.
  • float / int: a score. Flagged if a threshold is set and score >= threshold; with no threshold the score is only recorded.
  • a ProbeSignal: passed through untouched, for full control over action, confidence and metadata.

A flagged activation emits a signal with the probe's action; every other non-None return emits a CONTINUE signal, so scores reach signal_history and sinks. The verdict is {"flagged": bool, "max_score": float | None, "n": int}.

threshold and action can be overridden per extraction point through probe_args in the spec. Any other probe_args are passed to the function as keyword arguments, so declare them keyword-only. Function probes are single_shot, so they always run inline; write a Probe subclass for anything that keeps state across activations.

Parameters:

Name Type Description Default
name str | ProbeFunction | None

the probe_type to register under. Defaults to the function's __name__. @probe without parentheses also works.

None
threshold float | None

a numeric score >= threshold flags the activation. None (the default) records scores without flagging.

None
action str

what a flagged activation asks for: "abort" (default) stops generation at an inline extraction point, "flag" only records it.

'abort'
registry ProbeRegistry | None

where to register. Defaults to the global registry (default_registry).

None
register bool

set False to create the class without registering it.

True

Returns:

Type Description
Any

A FunctionProbe subclass named after the function. The function stays available as Cls.fn, for unit tests.

register_probe(name, probe=None, /, *, override=False, **probe_kwargs)

register_probe(name: str, /, *, override: bool = ..., **probe_kwargs: Any) -> Callable[[_P], _P]
register_probe(name: str, probe: type[Probe] | ProbeFactory, /, *, override: bool = ..., **probe_kwargs: Any) -> ProbeFactory

Register a probe under a probe_type name in the global registry.

Specs, ProbedModel and the Router then find it by that name.

@register_probe("linear_probe")          # decorator form
class LinearProbe(Probe):
    probe_kind = "single_shot"
    ...

register_probe("strict_linear", LinearProbe, threshold=0.9)  # functional form, default kwargs

Third-party packages can also make probes available without being imported, through an entry point in the undercurrent.probes group (name = probe_type, value = module:ProbeClass):

[project.entry-points."undercurrent.probes"]
my_probe = "my_package.probes:MyProbe"

Entry points are loaded the first time a lookup for their name misses. Explicit registrations always win over entry points.

Operates on default_registry; see ProbeRegistry.register for the arguments and errors.

FunctionProbe(*, threshold=..., action=..., **fn_kwargs)

Bases: Probe

Base class of every class @probe creates. Not used directly.

The decorator fills in the class attributes below. Instances hold the per-request state: the effective threshold/action, the extra kwargs for the function, and the signal history.

Attributes:

Name Type Description
fn ProbeFunction

the decorated function, unchanged.

probe_name str

the probe_type name the class was created with.

default_threshold float | None

the decorator's threshold.

default_action ProbeAction

the decorator's action, as a ProbeAction.

Probe lifecycle

Probe()

Bases: ABC

Base class for all probes.

Subclass it for stateful or trajectory probes (for a stateless function, @probe is shorter). Set probe_kind to "single_shot" or "trajectory" and implement the three lifecycle methods:

@register_probe("mean_norm")
class MeanNorm(Probe):
    probe_kind = "trajectory"

    def __init__(self, threshold: float = 10.0):
        super().__init__()
        self.threshold = threshold

    def on_start(self, request_ctx):
        self.norms = []

    def on_activation(self, record):
        self.norms.append(float(record.tensor.norm()))
        return None

    def on_end(self, request_ctx):
        mean = sum(self.norms) / max(len(self.norms), 1)
        return ProbeResult(self.request_id, self.extraction_point_name, verdict=mean > self.threshold)

Isolation: the subclass is a stateless, reusable factory. A fresh instance is spawned per (request, extraction point) with spawn, used for exactly one request (on_start -> on_activation* -> on_end) and then discarded. Keep per-request state on self, set in __init__ or on_start; a class-level list, dict, set or bytearray attribute is rejected when the class is defined, because every instance would share it.

Attributes:

Name Type Description
probe_kind str

"single_shot" (one activation in, one verdict out) or "trajectory" (state carried across many activations).

request_id property

The request this instance was spawned for.

extraction_point_name property

The extraction point this instance was spawned for.

spawn(request_id, extraction_point_name, **probe_kwargs) classmethod

Create a fresh probe instance scoped to one (request_id, extraction_point_name).

Every call returns a brand-new instance with its own state; nothing is shared with any other spawned instance, even of the same subclass constructed with identical kwargs. probe_kwargs are forwarded to the subclass's __init__ (e.g. a classification threshold), and are themselves fresh per call -- pass plain values, not shared mutable containers, if you want that guarantee to hold.

on_start(request_ctx) abstractmethod

Called once, before any activations, with the full request context.

on_activation(record) abstractmethod

Called once per matching activation.

May return a ProbeSignal immediately -- useful for inline probes that can intervene mid-generation (e.g. action=abort) -- or None if this activation doesn't warrant one.

on_end(request_ctx) abstractmethod

Called exactly once, when generation ends or is aborted.

Must always return a ProbeResult, even if on_activation was never called (e.g. the extraction point never matched a token).

ProbeFactory(probe_cls, probe_kwargs=dict()) dataclass

Binds a Probe subclass to fixed spawn kwargs.

A single value representing "the thing that produces probes for this extraction point", as stored in a probe registry or passed to Router(probe_registry={...}).

Attributes:

Name Type Description
probe_cls type[Probe]

the Probe subclass to spawn.

probe_kwargs dict[str, Any]

keyword arguments passed to its __init__.

spawn(request_id, extraction_point_name, **extra_kwargs)

Spawn a probe with probe_kwargs merged with extra_kwargs (extra_kwargs wins on a key clash -- the router passes an extraction point's probe_args here).

RequestContext(request_id, prompt_metadata, extraction_point_config) dataclass

What a probe knows about its request, passed to on_start and on_end.

Built by the router. Frozen: keep a probe's own mutable state on the probe instance (set in __init__ or on_start), not here.

Attributes:

Name Type Description
request_id str

identifier of the inference request.

prompt_metadata dict[str, Any]

router/adapter-defined metadata about the prompt (e.g. model name, prompt length, sampling params). Its shape is deliberately open-ended.

extraction_point_config Any

the resolved ExtractionPoint this probe was spawned for, so the probe can read its layer, tensor and position config.

Signals & results

ProbeSignal(action=ProbeAction.CONTINUE, metadata=dict(), confidence=None, timestamp=time.time()) dataclass

Returned by Probe.on_activation when an activation warrants a message.

Attributes:

Name Type Description
action ProbeAction

what the caller should do. Defaults to CONTINUE, i.e. "no intervention needed."

metadata dict[str, Any]

probe-defined, e.g. scores or intermediate values useful for logging/debugging.

confidence float | None

optional scalar confidence in this signal, meaning is probe-defined (e.g. classifier probability, distance from a threshold).

timestamp float

unix timestamp (seconds) of when this signal was created. Defaults to the time of construction.

ProbeAction

Bases: str, Enum

What the caller (router/adapter) should do in response to a signal.

CONTINUE = 'continue' class-attribute instance-attribute

No intervention needed.

ABORT = 'abort' class-attribute instance-attribute

Stop generation (honored for inline extraction points).

FLAG = 'flag' class-attribute instance-attribute

Record the activation as flagged; generation continues.

ProbeResult(request_id, extraction_point_name, verdict, signal_history=list(), metadata=dict()) dataclass

Returned by Probe.on_end, exactly once per (request, extraction point).

Attributes:

Name Type Description
request_id str

the request this result belongs to.

extraction_point_name str

which extraction point produced it.

verdict Any

the probe's conclusion. Deliberately untyped: a classifier probe might return a label + logits dict, a trajectory probe a running score, an intervention probe just a bool. Each probe should document its own verdict shape.

signal_history list[ProbeSignal]

every ProbeSignal this probe emitted during the request's lifecycle, in emission order. May be empty if the probe never intervened.

metadata dict[str, Any]

probe-defined auxiliary data that isn't part of the verdict itself (timing, debug info, etc.).

Registry

get_probe_factory(name)

Look name up in default_registry.

Raises:

Type Description
ProbeNotFoundError

nothing is registered under name.

list_probes()

Sorted names of every probe in default_registry, including plugins.

unregister_probe(name)

Remove name from default_registry.

ProbeRegistry(*, load_entry_points=True)

A thread-safe mapping of probe_type names to ProbeFactory.

Parameters:

Name Type Description Default
load_entry_points bool

whether lookup misses and list() consult installed plugins in the undercurrent.probes entry-point group. Defaults to True; pass False for a fully self-contained registry.

True

register(name, probe=None, /, *, override=False, **probe_kwargs)

register(name: str, /, *, override: bool = ..., **probe_kwargs: Any) -> Callable[[_P], _P]
register(name: str, probe: type[Probe] | ProbeFactory, /, *, override: bool = ..., **probe_kwargs: Any) -> ProbeFactory

Register a probe under name.

Used as a decorator (@registry.register("name")) it registers the decorated class and returns it unchanged. Called with a Probe subclass or a ProbeFactory it registers that and returns the stored factory. Extra keyword arguments become default spawn kwargs.

Raises ValueError if name is already registered (unless override=True, or the new class is a redefinition of the same class, e.g. a re-run notebook cell), and TypeError if the probe isn't a concrete Probe subclass with a valid probe_kind.

unregister(name)

Remove name. Raises ProbeNotFoundError if absent.

get(name)

Return the ProbeFactory registered under name.

On a miss, tries a matching undercurrent.probes entry point before raising ProbeNotFoundError.

list()

Sorted names of every registered probe, including all loadable plugins.

default_registry = ProbeRegistry() module-attribute

The process-wide registry behind the module-level functions and Router().

ProbeNotFoundError(name, known)

Bases: ProbingError, KeyError

Raised when no probe is registered under a requested probe_type.

Subclasses KeyError so mapping-style except KeyError handlers keep working. The message lists the known names, the closest matches and how to register a probe.

Example probes

undercurrent.core.examples holds two reference probes with deterministic stub scoring. They show the two probe shapes and are used in tests and demos; they are not production probes.

MLPClassifierProbe(num_classes=2)

Bases: Probe

single_shot probe: one activation in, one classification verdict out.

Uses a deterministic stub in place of a trained MLP head.

The verdict (set in on_end) is {"predicted_class": int, "logits": list[float]}.

Parameters:

Name Type Description Default
num_classes int

number of classes (logits) to produce.

2

TrajectoryScoreProbe(threshold=0.8)

Bases: Probe

trajectory probe: running-mean score with abort-on-threshold.

Scores each activation with a deterministic stub (the mean of its values) and emits an ABORT signal the first time the running mean reaches threshold.

The verdict (set in on_end) is {"final_mean": float, "count": int, "aborted": bool}.

Parameters:

Name Type Description Default
threshold float

running-mean score at which the probe asks to abort.

0.8

running_mean property

Mean score of the activations seen so far (0.0 before any).

check_intervention()

Called after each activation is folded into the running score.

Returns an ABORT signal the first time the running mean reaches threshold, and None otherwise (including every call after that, so the abort is emitted once).