simpleSTT
simpleSTT turns audio and video files into clean, accurate text using the best speech-to-text models — and sends your audio straight to them, with your own API key. Never an unknown server, never an unnamed model.
- Your choice of model — OpenAI (GPT-4o Transcribe, Whisper), Google Gemini and ElevenLabs Scribe; switch per job.
- Built for real work — M4A, MP3, WAV, FLAC, MP4, MOV and more; long recordings split automatically; language detection or a fixed language; custom prompts for vocabulary and formatting.
- Searchable history — every transcript stays on your device, with full-text search.
- Fair pricing — one free transcription a day; unlimited with a one-time purchase or a subscription. You pay your provider directly for what you use, usually cents per recording.
For iPhone and iPad.