Ashman Das

Grace

Building. Mid-way through moving from Python to Rust and Tauri. Entered in the IIT Delhi Youth Ideathon, September 2026.

Grace lets someone operate a Windows computer by voice, in their own words. It runs entirely on the machine.

  1. agentic_completion_rejected
  2. agentic_preexec_open_app
  3. agentic_read_pdf
  4. agentic_step_failed
  5. agentic_two_step
  6. agentic_unparseable_plan
  7. blind_app_vision_mode
  8. conversation_plain
  9. empty_transcript
  10. fastpath_adjust_volume
  11. fastpath_close_app
  12. fastpath_cua_list_windows
  13. fastpath_delete_file
  14. fastpath_lock_computer
  15. fastpath_open_app
  16. fastpath_open_calculator
  17. fastpath_open_file
  18. fastpath_search_files
  19. fastpath_tool_error
  20. followup_agentic
  21. followup_chain_three_turns
  22. rate_limited_intent
  23. rate_limited_planner
  24. repeated_action_escalates
  25. safety_confirm_accepted
  26. safety_confirm_declined
  27. safety_confirm_hijacked
  28. safety_press_key_confirm
  29. stale_field_replace
  30. unparseable_intent
All 30 recorded test runs. The 10 in red are named after a way Grace failed.

The problem

Voice control for people with motor disabilities still works like a command language. Say the exact phrase and it works. Say it another way and nothing happens. That is a second tax on the people who have the least to spare.

Grace drops the grammar. You say what you want, and it works out the steps.

How it works

A wake word starts it listening. Whisper turns speech into text and a local Gemma model understands it, both running on the laptop. A planner then works through Windows using the accessibility tree, which is the same structure screen readers use to know what is on the screen.

Some apps expose nothing to that tree. For those, Grace falls back to looking at the screen itself: it marks every clickable thing with a number and a vision model picks one.

Nothing leaves the machine. Someone who uses a computer by voice is dictating their whole life into a microphone.

Testing a loop with a microphone in it

Free speech is less predictable than a grammar. That unpredictability is the harm when the person can’t grab the mouse and undo a mistake. So anything that can’t be undone waits for a spoken yes.

To test it, Grace has 30 recorded scenarios. Each one replays the raw audio, the transcript and the model’s answers against a frozen clock, so the same input always produces the same run. Many are named after a way it actually failed: agentic_unparseable_plan, repeated_action_escalates, safety_confirm_hijacked.

Who it has to run for

The people Grace is for won’t buy a graphics card, and cloud APIs cost money every month. So it has to work on an ordinary Intel laptop with no GPU. The test machine for that is my old i5-8300H laptop. Faster hardware gets used when it’s there, but the slowest machine is the one that has to work.

Links