ReframeWeb is an experimental agentic workflow environment for replacing the legacy browser mental model.
It is not a conventional web browser with agents layered on top. Instead, ReframeWeb treats interactive work as a combination of agent-readable semantic capabilities, user-visible visual panels, persistent memory, and deterministic compute tools.
The project starts from the user interaction loop: Python receives audio, BAML drives the agentic flow, and the rest of the system is called from there. Native visual panels, a transport layer, WebAssembly Stores, Data Lenses, Compute Modules, and memory hang off that host-driven flow.
The legacy browser model is built around pages, tabs, browser chrome, DOM state, and visual interaction as the primary interface. ReframeWeb starts from a different assumption: agents and users should share a richer interaction surface where semantic capabilities are first-class.
In ReframeWeb, the first runtime surface is the Agent Host: a Python process that receives audio, coordinates BAML, and routes work through native windows, transport calls, Stores, Lenses, Compute Modules, and memory.
The shared substrate for those components is a typed local protocol. A Semantic Store is an installed WebAssembly Component that exposes resources, functions, schemas, and focused usage guidance through progressive discovery rather than existing as a normal application library. The reusable Rust host and conformance CLI live in semantic-store.
CEF is used as a rendering and native windowing foundation for Visual Panels. It is not the product's central abstraction. The goal is to build a new agent-native workflow layer, not to preserve the browser as the underlying concept.
- Not a browser extension.
- Not Selenium-style browser automation.
- Not a legacy browser with an assistant bolted onto it.
- Not a DOM-first control system where agents primarily click their way through pages.
- Agent Host: A Python manager that handles audio input, transcription, typed data and side-effect boundaries, memory persistence, and TTS playback.
- BAML Flow: The complete routed voice-turn orchestrator. Audio is transcribed and handed to one top-level BAML flow that owns agentic decisions, branches, retries, and task reselection.
- Transport Layer: Local typed routing, framing, calls, streaming, and cancellation. The Semantic Store protocol and platform IPC are implemented; integration with Lenses, Views, and Compute remains future work.
- Semantic Store: A locally hosted WebAssembly Component whose typed resources and functions agents can progressively discover, inspect, and call.
- Visual Panel: A React display surface shown in a native CEF-backed panel window. Users can see state and occasionally click or scroll, but required functionality should be exposed for agent-driven control.
- Data Lens: An optional Rust support layer between a Store and a Visual Panel. It is for cases where behavior needs to differ substantially from the original site or application intent, or where the Store lacks a sensible API surface for React to parse directly.
- Compute Module: A short Rust script for specialized, complex, repetitive deterministic tasks. Compute Modules are not the default path for ordinary repeated behavior, because overusing them would slow down the agentic flow.
- Memory Graph: A graph-backed memory system for preferences, task context, user habits, and recent relevant memories.
- Agent Workspace: A transparent projected filesystem for Codex, OpenCode, and their child processes. It combines memory-derived files, ephemeral working state, direct-disk dependency trees, and policy-controlled retained snapshots.
ReframeWeb is driven by a small set of deliberate technology choices:
More detail is tracked in Technology.
The projected agent workspace is specified separately in Agent Workspace Architecture Decisions, with a phased delivery plan in the Agent Workspace Implementation Design.
- Rust for native CEF/window agentic bindings, the projected workspace daemon, Data Lenses, Compute Modules, memory-related runtime components, and likely Store implementations compiled to WebAssembly.
- CEF as the embedded rendering and native windowing foundation for Visual Panels. CEF is infrastructure here, not the conceptual model of the product.
- React for Visual Panel content, display state, Store-backed data fetching, and any user-visible controls.
- Python for the Agent Host, which owns setup, audio processing, transport and memory I/O, persistence, and TTS playback through typed BAML boundaries.
- BAML for the complete routed voice-turn flow and all agentic decisions, branches, loops, retries, and task reselection.
- WebAssembly Components for portable Semantic Stores, hosted by an asynchronous Rust/Wasmtime runtime with WASI HTTP.
- Graph database storage for memory, including tagged memory nodes, descriptions, created/read/modified timestamps, relationships, and recency-aware retrieval.
sounddevicefor microphone input and audio playback control.pocketsphinxfor local keyphrase spotting, including commands such as "jarvis do x" and a "conversation on" mode trigger.silero-vadfor voice activity detection.faster-whisperandwhisper.cppfor transcribing spoken prompts before they are passed into the BAML-driven agentic flow.kokoro-onnxfor TTS playback, using theaf_heartvoice.
The audio layer should support cancelling spoken playback when the user starts talking without destroying the underlying work the agent was already performing.
This repository now contains a working Agent Host prototype. The Python voice loop, graph-backed memory flow, the cross-platform Agent Workspace, and the presentation-independent Semantic Store host are active code. The CEF/React Visual Panel layer, Data Lens runtime, Compute Module runtime, and integration between these subsystems remain future-facing architecture.
Implemented pieces include:
- A
uv-managed Python Agent Host with CLI commands for setup, checks, voice turns, memory seeding, and benchmarks. - Local wake/phrase detection, VAD, local Whisper transcription, and turn recording into the memory graph.
- BAML stages for task choice, conversation memory-search hint generation, and per-domain search-depth selection.
- A SurrealDB-backed memory graph with roots for providers, tasks, sessions, conversations, session memories, task-choice memories, conversation evaluation memories, and search-depth memories.
- Graph-based memory retrieval that starts from existing roots and relations, applies search hints and timestamp breadth to candidate nodes, and hydrates parent wrappers needed to explain valid child matches.
- Benchmark harnesses for task choice, conversation evaluation, and control flow/search-depth behavior.
- A Rust Agent Workspace daemon and uv-integrated Agent Host surface with
Windows ProjFS plus Linux/macOS FUSE mounts, resident deduplicated file bytes,
SQLite metadata, BLAKE3 retained blobs, SurrealDB-backed filesystem memory
nodes, native scratch paths, journals, checkpoints, and remount support.
The daemon is installed by
uv syncand controlled through the existingreframe-agent-host workspaceCLI over user-local IPC. See Agent Workspace usage. - A cross-platform Rust Semantic Store workspace with a strict typed package format, bounded progressive catalog, isolated Wasmtime Component execution, WASI HTTP, streaming and cancellation, secure user-local IPC, a guest SDK, reference Store, and conformance CLI. See the Semantic Store host guide.
On Windows, enable Git's long-path support once before resolving the pinned BAML fork. The fork contains deeply nested compiler snapshots that exceed Git's legacy 260-character checkout limit:
git config --global core.longpaths truecd agent-host
uv sync
uv run reframe-generate-baml
uv run reframe-check-baml
uv run reframe-agent-host doctorUse reframe-generate-baml after editing agent-host/baml_src. It checks the
BAML project, runs the deterministic workspace BAML tests, generates the Python
SDK, and normalizes embedded source paths so generated artifacts are portable
between checkouts. reframe-check-baml runs the same pipeline and fails if it
changes a generated file; use it in CI and before committing. Both commands
resolve the developer checkout from the current directory and fail clearly
outside one; BAML sources are not part of the runtime wheel. Do not invoke
baml generate directly, because its raw bytecode module contains
checkout-specific absolute source paths.
The Agent Host uses OpenCode Go through its OpenAI-compatible endpoint and reads
the API key from OPENCODE_GO_API_KEY.
Model use is a BAML choice, not user-selected global configuration and not
Python-side model routing. Each agentic task should have an explicit model
assignment through a memory Provider node that points at a BAML surface. The
current task-choice, conversation-evaluation, and search-depth flows use
glm-5.1 without reasoning effort.
Current benchmarked OpenCode Go model IDs include:
kimi-k2.7-codekimi-k2.6kimi-k2.5glm-5.1glm-5deepseek-v4-prodeepseek-v4-flashmimo-v2.5-promimo-v2.5
The first runnable microphone path is intentionally narrow and testable:
cd agent-host
uv run reframe-agent-host transcription-check
uv run reframe-agent-host audio-devices
uv run reframe-agent-host voice-turn --device 1The project config sets uv's package install link mode to copy, which avoids
Windows hardlink warnings when uv's cache and this repository live on different
drives.
If uv is not on PATH on Windows, use the local runner instead:
cd agent-host
.\reframe-agent-host.cmd transcription-check
.\reframe-agent-host.cmd audio-devices
.\reframe-agent-host.cmd voice-turn --device 1voice-turn listens for one utterance, detects the speech boundary, transcribes
it with the configured local Whisper backend, records the current turn when
session and conversation IDs are available, sends the transcript through the
BAML control flow, retrieves graph memory context, and prints concise per-stage
summaries and latencies. When memory retrieval runs, the CLI prints the
retrieved memories directly instead of dumping the full turn result as JSON.
The current top-level BAML voice-turn path is:
- Choose an initial task from the task catalog.
- Generate memory search hints from the conversation and selected task.
- Choose timestamp breadth for each search domain.
- Retrieve graph memories through a typed host boundary.
- Select relevant memories and compose the task prompt.
- Execute the task through a typed host boundary and summarize its actions.
- Run the existing completion pass/fail check.
- On failure, generate and persist a
validation_reply, then either retry the same task with stacked task-session replies or return to task selection. - On success, return the turn result without retaining task-local retry data.
Python supplies data and performs external side effects at the named boundaries. It does not decide which BAML path runs. The complete agentic turn is kept in a single BAML graph so the operational flow remains inspectable.
Memory retrieval is relation-first rather than table-wide. The task catalog is
searched from the task root. Past conversation context is searched through
session, conversation, message, and session-memory relations. Search hints are
alternatives, so any positive tag or string hint can match a candidate. Timestamp
breadth is restrictive: candidate nodes must pass created_at, updated_at,
and, when present, read_at cutoffs. Parent wrappers are included when a valid
child match needs them for context. Current-session memories are always included
in the retrieved memory output for the active session.
By default, WakeCommand mode is wake-gated locally with a rolling PocketSphinx
phrase recognizer. It does not require an account, network call, or paid
wake-word service. Say the single-word trigger "jarvis" followed by the prompt.
The phrase "conversation on" switches the host into continuous conversation
mode. If more speech follows in the same utterance, such as "conversation on
this is a test", the trigger audio is trimmed away and the remaining command
audio is sent through VAD and local transcription.
uv run reframe-agent-host voice-turn --device 1 --no-task-choiceOn Windows, Linux, and macOS, the default transcription backend tries CUDA
faster-whisper first, then whisper.cpp, then CPU faster-whisper. CUDA is
still supported when available. Non-CUDA GPU support comes through a
whisper.cpp binary compiled for the target backend, such as Metal/Core ML on
macOS, Vulkan or OpenVINO on Windows/Linux, and ROCm on Linux.
Examples:
uv run reframe-agent-host transcription-check --transcriber whisper-cpp --transcriber-device metal --whisper-cpp-bin /path/to/whisper-cli --whisper-cpp-model /path/to/ggml-model.bin
uv run reframe-agent-host voice-turn --transcriber whisper-cpp --transcriber-device vulkan --whisper-cpp-model /path/to/ggml-model.bin
uv run reframe-agent-host voice-turn --transcriber faster-whisper --transcriber-device cpu --whisper-cpu-compute-type int8Windows runner equivalent:
.\reframe-agent-host.cmd voice-turn --device 1 --no-task-choiceUseful tuning flags:
--wake-keyword jarvisto configure local wake keyphrases.--conversation-on-phrase "conversation on"to configure the local conversation-mode trigger. The phrase is treated as control input only, not as a user request.--conversation-on-confirm-window-ms 2000to tune how much recent audio is sent to Whisper for the second-layer confirmation.--wake-gain 1.0to tune local phrase detector gain without changing the audio sent to Whisper.--wake-threshold 1e-30to tune local PocketSphinx KWS confirmation for wake-keyword candidates.--wake-replay-pre-ms 0starts command replay at the confirmed wake boundary so Whisper does not need to transcribe the wake word.--debug-audio-dir .debug-audioto opt in to saving local WAV clips and JSON sidecars for missed wake timeouts and detected keyphrases.--debug-audio-seconds 8to tune how much rolling microphone audio is kept for those debug clips.--debug-audio-period-seconds 5to opt in to saving rolling clips every few seconds while waiting for a wake phrase.--vad sileroto require Silero VAD, or--vad energyto use the simple RMS fallback.--vad-threshold 0.35to tune Silero sensitivity for quieter microphone input.--min-silence-ms 0is the default. Increase it only if the utterance cuts off too early.--final-silence-ms 1450controls how long a provisional endpoint can be cancelled if speech resumes after Whisper starts.--pre-speech-ms 320keeps audio just before VAD start so first syllables are not cut off.--wake-carry-ms 220keeps audio around wake detection so commands that start immediately after the wake word are not clipped.--energy-start-threshold 0.02if the fallback detector starts too easily.--transcriber faster-whisperor--transcriber whisper-cppto choose a backend explicitly.--transcriber-device cuda,cpu,metal,coreml,vulkan,openvino, orrocmto choose the preferred local runtime.--whisper-model large-v3to choose a faster-whisper model name or local faster-whisper model path.--whisper-cpp-model /path/to/ggml-model.binto use a local whisper.cpp ggml model.--whisper-cpp-bin /path/to/whisper-cliwhen the whisper.cpp binary is not on PATH.--whisper-compute-type int8_float16to test lower CUDA memory use than the defaultfloat16.--whisper-cpu-compute-type int8to choose the CPU fallback compute type.--no-cpu-fallbackto fail fast when the requested GPU backend is unavailable.
On Windows, the CUDA faster-whisper path checks for CUDA 12 cuBLAS DLLs before
listening, including project-local agent-host\.cuda\bin, normal CUDA Toolkit
installs, and NVIDIA Python wheel locations such as nvidia-cublas-cu12.
Saved debug WAVs can be replayed through the local wake recognizer without calling Whisper:
uv run reframe-agent-host debug-wake-audio ".debug-audio/*.wav"ReframeWeb is being designed around a few working principles:
- Agent-native semantic interfaces should be the primary interaction layer.
- The first implementation path starts where the user interacts: audio into the Python Agent Host, then BAML-driven agentic flow.
- Visual Panels are React display outputs for showing useful state and controls. Users may click or scroll when they want to, but core functionality should be available to the agent rather than requiring manual UI operation.
- Compute Modules are for specialized, complex, repetitive deterministic behavior where a module genuinely improves reliability or clarity. Ordinary repeated actions should stay in the Store and agent flow.
- Personalization should usually happen through Visual Panels and Store-backed presentation. Data Lenses are for larger behavioral changes or for supporting cases where the Store does not expose a sensible API surface for React.
- Incremental development should build the actual project architecture rather than a substantially cut-down or throwaway version of it.