Voice-Trajectory Spatiotemporal Compiler

Point, Circle & Speak
Compile Motions into AI Action

In-Motion Voice Binding · Word-Level Timestamp Alignment · Codex MCP · Chrome Extension

Typing cannot convey spatial coordinates, while static screenshots fail to communicate fluid intent. iMouse allows you to speak naturally while pointing, circling, and drawing lines—compiling trails and voice into structured evidence packages in milliseconds to drive AI code, design, and web tasks.

In-Motion Voice-Location BindingStructured Instruction CompilerCodex Native MCP PluginChrome Web Extension Dual-Mode

Trajectory-Voice Alignment · Structured Instruction Compilation · Codex MCP Plugin · Offline SenseVoice ASR

design-review.png

"Make this button lighter… and move this whole block to the left"

Compiled · Structured Evidence Pack
{ "action": "adjust_color",
  "target": { "shape": "circle" },
  "hint": "lighter" }
{ "action": "move_block",
  "target": { "arrow": "→ left" } }

Core Architecture & Capabilities

Engineered for high-precision human-AI visual collaboration

L1

In-Motion Voice-Location Binding

Never stop the mouse to type. Speak continuously while dragging, circling, or underlining, binding speech to spatial trajectory with millisecond precision.

L2

Structured Instruction Compiler

Built-in geometric shape recognition (enclosures, arrows, underlines) and dwell-time heatmap algorithms compiling raw trails into contextual evidence packages.

L3

Codex Native MCP Plugin Loop

Seamlessly integrates with Codex via MCP to automatically pull pending visual tasks, viewport screenshots, and trajectory metadata for 2-way execution.

Chrome Browser Extension Dual-Mode

Interact on any live webpage with Alt + Drag. Supports both standalone AI sidebar assistant mode and Codex collaborative workflow.

3-Tier ASR Speech Recognition Matrix

Prioritizes OpenAI word-level timestamp endpoints, with local offline SenseVoice int8 model (zero latency, 100% private) and Apple Speech fallback.

Layered Architecture & Zero Native Dependency

Clean decoupled design across Pointer Capture → Instruction Compiler → Agent Adapter layers, ensuring ultra-fast web and extension performance.

Why We Built iMouse?

Traditional AI interfaces are constrained to text boxes. When looking at UI mockups, webpages, or code editors and trying to tell AI 'move this block over there and lighten this button', users have to repeatedly switch windows, stop mouse movement, and type clumsy descriptions.

The most natural human collaboration is 'pointing with a finger while speaking'. iMouse captures pointer trajectories and performs spatiotemporal fusion with SenseVoice/Whisper word-level timestamps. It compiles human gestures and voice into deterministic, actionable agent instructions, delivering true What-You-Point-Is-What-You-Get interaction.

Unite gesture and voice: What you point and say is what you get.