SKI: Free Local Voice Coding for Claude Code, Codex & OpenClaw

Talk to Claude Code, Codex, Cursor, and other coding agents through local speech recognition and on-device voice replies.

SKI is a free desktop app for Mac and Windows that adds a spoken conversation loop to your AI coding agent.

It converts speech to text on the device, sends the request to Claude Code, Codex, Cursor, Gemini CLI, Windsurf, or OpenClaw, then speaks the agent’s reply through a local neural voice.

The app also includes a local meeting recorder that captures a call’s audio and transcribes it on-device, and an optional paid service, AgentCall, that sends the connected agent into a video call as a participant.

What SKI Adds to an AI Coding Agent

SKI creates a two-way voice layer around an existing coding session.

A spoken request passes through local speech recognition and reaches the connected agent as text. The agent works inside its normal project environment and sends a response back to SKI. A local text-to-speech model reads the selected response aloud while the desktop pill displays the text and current agent status.

The agent controls the spoken message. Large code blocks, test logs, and full diffs stay on screen. Spoken output covers a brief status update, completion summary, error explanation, or request for additional input.

SKI App Speek Agent

Where SKI Fits Beside Native Voice Features

Claude Code and Codex include their own voice options.

Claude Code’s /voice command records a prompt and sends the audio for transcription. The recognized text appears in the terminal input. It does not create the same cross-agent desktop voice layer offered by SKI.

The ChatGPT desktop app supports live voice interaction with Codex for eligible accounts. Users speak, interrupt responses, and coordinate coding tasks through Codex.

SKI focuses on one desktop voice across several coding agents, local speech processing, multiple connected repositories, project-specific voices, screenshots, and meeting tools.

CapabilitySKIClaude Code Voice InputCodex Voice
Voice inputLocal speech recognitionBuilt-in transcriptionChatGPT Voice
Spoken repliesYesNo full spoken reply loopYes
Agent coverageMultiple supported agentsClaude CodeCodex
Multiple projectsConcurrent project connectionsCurrent Claude Code sessionsCodex projects and folders
Project-specific voicesYesNoNo
Screenshot contextHotkey or agent requestSeparate workflowDesktop permissions
Local meeting recorderIncludedNoNo
AI meeting participantAgentCall cloud featureNoNo

Key Features

  • Local speech input: Speech recognition runs on the computer and passes text to the connected coding agent.
  • Local spoken replies: A neural voice reads short agent responses, status messages, and requests for input.
  • Full-duplex interruption: SKI detects new speech while a reply is playing and stops to accept the next instruction.
  • Multiple project connections: Several repositories stay connected at the same time. Each project receives its own voice.
  • Agent status indicator: A green status light shows when the selected coding agent is connected and listening.
  • Editable transcript approval: Review mode displays the recognized prompt in an editable bubble before delivery.
  • Screenshot context: A global hotkey captures the screen and attaches it to the next spoken request.
  • Silent mode: Replies appear as text in the pill while speech output stays muted.
  • Local meeting transcription: SKI records microphone and system audio on separate tracks, generates a live transcript, and exports Markdown or plain text.
  • AgentCall meeting participation: A connected agent joins Google Meet, Microsoft Teams, or Zoom as an active participant or silent notetaker.
SKI App Running on Mac

Voice Coding Workflows

Direct a Coding Task Away From the Keyboard

Open a connected repository and describe the result aloud:

Add validation to the signup form, cover the new rules with tests, and tell me when the test suite passes.

The coding agent receives the transcript, inspects the project, changes the files, and runs its normal tools. SKI speaks a brief result after the task completes or asks for clarification when the agent reaches a decision point.

Discuss a Test Failure

A long test run finishes while another window has your attention. Ask SKI which tests failed and what caused the failure. The agent reads the test output and speaks a compact explanation while the complete logs stay available in the coding interface.

Show a Visual Problem

Capture the current screen with the screenshot hotkey, then describe the issue:

The mobile menu overlaps the page title. Fix the layout shown in this screenshot.

The image travels with the spoken prompt. This workflow suits visual bugs, browser errors, terminal output, and interface details that take too many words to describe precisely.

Move Between Several Repositories

Connect a frontend project, an API service, and a documentation repository. Select the target project from the SKI pill before speaking. Separate voices identify which project is replying when several agent sessions stay active.

Turn a Meeting Into Project Notes

Start the local meeting recorder before a planning call. SKI captures the microphone and system audio, writes a timestamped transcript, and stores it on the computer.

After the call, open the transcript and ask the connected coding agent to extract technical decisions, unresolved questions, and action items.

Local Recording and AgentCall Meetings

SKI offers two meeting workflows with different privacy and pricing boundaries.

Local Meeting Recorder

The local recorder captures the user’s microphone and system audio as two labeled tracks: “You” and “Meeting.” It works with any meeting or webinar that plays audio through the computer.

Recording time and local retention have no usage cap. Transcripts stay on the device and export as Markdown or plain text. Individual remote speakers do not receive separate identity labels in this mode.

AgentCall Meeting Participant

AgentCall places the coding agent inside a video call. The agent joins as an active participant or silent notetaker, listens to individual speakers, reads chat, views shared screens, sends messages, and presents its own screen.

This mode runs in the cloud. Meeting audio, transcripts, chat, and related data pass through AgentCall infrastructure. Usage is billed per minute after the included free allowance.

Meeting recording and AI participation are subject to local consent, privacy, and wiretapping laws. The user is responsible for obtaining required permission from participants and following each meeting platform’s terms.

SKI Privacy and Data Protection

SKI’s local features keep sensitive development data on the computer, but the entire product does not operate under a zero-network model.

Data or FeatureProcessing LocationDetails
Spoken audioLocal deviceSKI does not upload audio from the local voice loop
Speech transcriptionLocal deviceOn-device speech recognition
Synthesized voiceLocal deviceOn-device neural speech
SKI transcriptsLocal deviceStored until the user deletes them
Code and project pathsLocal device within SKIThe connected coding agent follows its own data policy
Local meeting recordingsLocal deviceView, export, or delete from the Transcripts window
Product analyticsExternal analytics serviceAnonymous usage statistics are configurable in Preferences
Account registrationAgentCallA free sign-in is required after the trial
AgentCall meetingsCloudProcesses meeting audio, transcript, chat, and related data

How to Use SKI

1. Download the DMG for an Apple Silicon Mac or the EXE installer for Windows.

2. On macOS, drag SKI into the Applications folder and approve the first-launch security prompt. On Windows, run the installer.

3. Choose the microphone, local voice, and activation hotkey in the setup wizard.

4. Open Claude Code, Codex, Cursor, Gemini CLI, Windsurf, or OpenClaw and install the corresponding SKI skill or project connection.

ski

4. Start with a simple request in your coding agent. The agent performs the task through its existing tools and returns the result through SKI.

5. Configure Spoken Responses. This prompt keeps code, diffs, and long explanations on screen while reserving audio for useful status changes.

Speak only when a task finishes, a test fails, or you need a decision from me.

6. Open Preferences → Transcription → Approve before send to enable Approval for Sensitive Work.

Pros

  • Local speech recognition
  • Local neural voice output
  • Multiple connected projects
  • Spoken agent status updates
  • Editable transcript approval
  • Screenshot attachments
  • Free local meeting transcription
  • Markdown and text exports
  • Configurable global hotkeys

Cons

  • English only
  • Free account required
  • Paid cloud meeting participation
  • Transcript-only action review

Alternatives & Related Resources

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the latest & top AI tools sent directly to your email.

Subscribe now to explore the latest & top AI tools and resources, all in one convenient newsletter. No spam, we promise!