DocsPerception

DocsTechnology

Perception

Perception turns an observed window into structured context for an agent. Read the current controls, text and layout before choosing a target.

Read a Scene

A Scene describes what Mecum identifies in a window: text, controls, states, groups, panels and positions. It is an observation of the current interface, not a complete inventory of the app’s capabilities.

SignalContribution
OCRVisible text: labels, values, menu items. Apple’s text recognition, tuned for interface labels; tiles that did not change are not read again.
Visual candidatesEdge-based regions for icons and controls without readable text. Pixels alone give a candidate and its bounds, never a name, a role or whether it can be acted on.
PanelsColor-based sections that group elements into named areas, such as a sidebar or a toolbar.
Accessibility, when availableRoles, values and state the app exposes. It only adds, within a time budget, and only where its frames agree with the window server.
Control state from pixelsA switch’s knob side or a checkbox mark, filled in only where the app said nothing.

A Scene contains text and window-relative positions, with no image data.

Window context as text

The model reads the Scene as lines: the app and window, the viewport, then each element with its kind, label, state, group and position, nested under its panel.

Illustrative Scene · invented values
app: TextEdit (com.apple.TextEdit): "Project notes"
viewport: 1512x869
elements (4) in 2 sections:
Section: toolbar, position: 0.00,0.00 1.00×0.08, 2 elements
    [control] Bold [off] (Formatting)  @ 0.12,0.03
    [control] Helvetica = 12 (Formatting)  @ 0.30,0.03
Section: document, position: 0.00,0.08 1.00×0.92, 2 elements
    [text] Project notes  @ 0.05,0.12
    [icon?] id:'e7'  @ 0.92,0.10

An unlabeled icon is still a target: its ID can be acted on, and the Brain can name it later.

Resolve the live target

Window references and positions belong to an observation. A moved window, a changed layout or a newly opened dialog can invalidate them. Each observation carries a revision; an action must follow a current one, and the runtime observes again before acting.

  1. Identify the intended app and window.
  2. Read the current Scene.
  3. Resolve the control using its label, role and surrounding group.
  4. Check the live target before acting, then observe the result.

A remembered position cannot replace these checks.

Handle ambiguity

SituationResponse
Repeated labelsThe tool answers ambiguous. Use the containing panel, the role and the current state, or the element ID, to narrow the target.
Missing or uncertain textThe tool answers honest_miss. Request a fresh observation of the relevant window, or use the element ID.
Dialog appearsResolve the dialog as a target; do not reuse the parent window’s geometry.
Accessibility disagrees with pixelsTreat the mismatch as uncertain context. Inspect before choosing a target.

If the target remains ambiguous, ask for clarification or stop. Do not choose an arbitrary matching control.

Local processing, model context

The window-to-text pipeline runs on the Mac. The resulting text can still leave the Mac when included in a request to a remote model.

Permissions & data explains access and data boundaries. Architecture shows where observation, model reasoning and execution connect.

Verified source

Checked against Mecum app source at main 524eb7f on October 1, 2026. UI labels and paths have not yet been checked against the signed release build.