Point, Circle & Speak
Compile Motions into AI Action
In-Motion Voice Binding · Word-Level Timestamp Alignment · Codex MCP · Chrome Extension
Typing cannot convey spatial coordinates, while static screenshots fail to communicate fluid intent. iMouse allows you to speak naturally while pointing, circling, and drawing lines—compiling trails and voice into structured evidence packages in milliseconds to drive AI code, design, and web tasks.
Trajectory-Voice Alignment · Structured Instruction Compilation · Codex MCP Plugin · Offline SenseVoice ASR
"Make this button lighter… and move this whole block to the left"
{ "action": "adjust_color",
"target": { "shape": "circle" },
"hint": "lighter" }
{ "action": "move_block",
"target": { "arrow": "→ left" } }Core Architecture & Capabilities
Engineered for high-precision human-AI visual collaboration
In-Motion Voice-Location Binding
Never stop the mouse to type. Speak continuously while dragging, circling, or underlining, binding speech to spatial trajectory with millisecond precision.
Structured Instruction Compiler
Built-in geometric shape recognition (enclosures, arrows, underlines) and dwell-time heatmap algorithms compiling raw trails into contextual evidence packages.
Codex Native MCP Plugin Loop
Seamlessly integrates with Codex via MCP to automatically pull pending visual tasks, viewport screenshots, and trajectory metadata for 2-way execution.
Chrome Browser Extension Dual-Mode
Interact on any live webpage with Alt + Drag. Supports both standalone AI sidebar assistant mode and Codex collaborative workflow.
3-Tier ASR Speech Recognition Matrix
Prioritizes OpenAI word-level timestamp endpoints, with local offline SenseVoice int8 model (zero latency, 100% private) and Apple Speech fallback.
Layered Architecture & Zero Native Dependency
Clean decoupled design across Pointer Capture → Instruction Compiler → Agent Adapter layers, ensuring ultra-fast web and extension performance.
Why We Built iMouse?
Traditional AI interfaces are constrained to text boxes. When looking at UI mockups, webpages, or code editors and trying to tell AI 'move this block over there and lighten this button', users have to repeatedly switch windows, stop mouse movement, and type clumsy descriptions.
The most natural human collaboration is 'pointing with a finger while speaking'. iMouse captures pointer trajectories and performs spatiotemporal fusion with SenseVoice/Whisper word-level timestamps. It compiles human gestures and voice into deterministic, actionable agent instructions, delivering true What-You-Point-Is-What-You-Get interaction.
Unite gesture and voice: What you point and say is what you get.