Grace
Building. Mid-way through moving from Python to Rust and Tauri. Entered in the IIT Delhi Youth Ideathon, September 2026.
Grace lets someone operate a Windows computer by voice, in their own words. It runs entirely on the machine.
- agentic_completion_rejected
- agentic_preexec_open_app
- agentic_read_pdf
- agentic_step_failed
- agentic_two_step
- agentic_unparseable_plan
- blind_app_vision_mode
- conversation_plain
- empty_transcript
- fastpath_adjust_volume
- fastpath_close_app
- fastpath_cua_list_windows
- fastpath_delete_file
- fastpath_lock_computer
- fastpath_open_app
- fastpath_open_calculator
- fastpath_open_file
- fastpath_search_files
- fastpath_tool_error
- followup_agentic
- followup_chain_three_turns
- rate_limited_intent
- rate_limited_planner
- repeated_action_escalates
- safety_confirm_accepted
- safety_confirm_declined
- safety_confirm_hijacked
- safety_press_key_confirm
- stale_field_replace
- unparseable_intent
The problem
Voice control for people with motor disabilities still works like a command language. Say the exact phrase and it works. Say it another way and nothing happens. That is a second tax on the people who have the least to spare.
Grace drops the grammar. You say what you want, and it works out the steps.
How it works
A wake word starts it listening. Whisper turns speech into text and a local Gemma model understands it, both running on the laptop. A planner then works through Windows using the accessibility tree, which is the same structure screen readers use to know what is on the screen.
Some apps expose nothing to that tree. For those, Grace falls back to looking at the screen itself: it marks every clickable thing with a number and a vision model picks one.
Nothing leaves the machine. Someone who uses a computer by voice is dictating their whole life into a microphone.
Testing a loop with a microphone in it
Free speech is less predictable than a grammar. That unpredictability is the harm when the person can’t grab the mouse and undo a mistake. So anything that can’t be undone waits for a spoken yes.
To test it, Grace has 30 recorded scenarios. Each one replays the raw audio, the transcript and the model’s answers against a frozen clock, so the same input always produces the same run. Many are named after a way it actually failed: agentic_unparseable_plan, repeated_action_escalates, safety_confirm_hijacked.
Who it has to run for
The people Grace is for won’t buy a graphics card, and cloud APIs cost money every month. So it has to work on an ordinary Intel laptop with no GPU. The test machine for that is my old i5-8300H laptop. Faster hardware gets used when it’s there, but the slowest machine is the one that has to work.