What cloud dictation actually sends
Dictation is not a keyboard. It captures a recording of your voice and, in the cloud case, uploads it. Whatever you were saying at the time goes with it: the client name, the salary figure, the half-formed idea you would never type into a web form.
Vendors handle that responsibly or they do not, and you cannot check from the outside. Reading a retention policy tells you what a company promises today. On-device processing removes the question instead of answering it, which is why I stopped weighing policies and changed tools.
Spokenly, free and on-device
Spokenly runs local speech models on Apple Silicon and types into any text field. Hold a shortcut, talk, release, and the text appears in whatever app has focus. Nothing goes out over the network, so it also works with the Wi-Fi off, on a plane, in a hotel with captive portal nonsense.

Text recognised inside a copied image, then searched from the same window. Recognition happens on the Mac, so the image is never uploaded to run it.
The pattern is the same one that makes local dictation worth the trade. Processing on device removes the question of what a server keeps, rather than answering it with a policy page.
It can also call cloud models if you point it at them, which puts you back where you started. I leave that switched off. The local path is the reason I installed it.
The honest limitation: local models have to be downloaded, the first run takes patience, and transcription quality trails the best cloud services on long dictation and on heavy technical vocabulary. Punctuation and formatting are also less clever than a paid service that post-processes your text with a language model.
