Skip to content

AI correction

AI polishes your raw transcription — from spacing, punctuation, and misrecognized-word fixes to filler removal and structuring. Choose a provider and a correction mode in Settings → LLM. The default is None (original text as-is), so you only turn correction on when you need it.

ProviderLocationVisionRequirements
None (use original) (default)——None. Inserts the transcription as-is
Local MLXLocal (Apple Silicon)Depends on modelModel download. Some MoE models need uv
OpenAI (GPT)Cloud✅Codex CLI token or OpenAI login
Groq CloudCloudLlama 4 Scout onlyGroq API key (shared with STT)
Claude (subscription)Via local claude CLI✅Claude Code CLI installed + logged in

Fully local correction. It supports text models (Qwen3, Gemma 4, GLM, etc.) and a vision model (Qwen3-VL-4B); download models in Settings → Models. With a vision model, you can leverage visual context for correction. For per-model compatibility, see Models & compatibility.

Uses the ChatGPT Responses API (SSE streaming). For authentication, it prefers reusing the Codex CLI token (~/.codex/auth.json), and if absent, connects via OpenAI login (browser PKCE) in the LLM tab. Models: GPT-5.5 (default) · 5.4 · 5.4 Mini · 5.3 Codex · 5.2. Vision supported.

OpenAI-compatible cloud. Models: Qwen3 32B (default) · Llama 3.3 70B · Llama 3.1 8B · GPT-OSS 120B/20B, with vision supported on Llama 4 Scout only. Uses the same Groq key as STT.

Calls the local claude -p CLI to reuse your Claude subscription directly (the only subscription path allowed under the ToS). Models: Haiku 4.5 (default) · Sonnet 4.6 · Opus 4.8 — all support vision. Responses usually take 5–20 seconds (including cold start), and billing draws from your subscription credit pool.

These appear in Settings → LLM when the provider isn’t ‘None’. The prompts are Korean-based and include Korean↔English code-switching examples.

ModeWhat it does
Standard (default)Fixes only obvious STT errors — spacing, punctuation, misrecognized words.
Filler removalStandard + removes verbal fillers like “um/uh/you know” (keeps the content).
StructuredStandard + filler removal + tidies rambling speech into bullets/numbers (not a summary, just organization).
CustomUses a system prompt you write yourself.

In Custom mode, you can edit the system prompt directly. In other modes, you can preview the currently applied prompt read-only, so you can see what instructions the correction follows.

Even when English technical terms get mixed into Korean speech (e.g., “밸리데이션” → “validation”), it corrects them appropriately. The code-switching prompt is enabled when the Language setting is auto/Korean.