Press Win + H on any Windows 11 machine and a microphone overlay appears. Start speaking and your words flow into whatever app is open — Outlook, Word, a browser text field, anywhere. It's fast, accurate, and deeply integrated into the OS. It's also sending your audio to Microsoft by default.
This post covers what the Windows Voice Typing privacy settings actually control, what Microsoft's documentation says about data handling, and where the gaps are for professional users.
Two settings, one dictation feature
Windows Voice Typing involves two separate privacy settings that are easy to confuse:
- Voice Typing. The feature itself (Win + H shortcut). Enabled by default. When active, audio is processed — but where it's processed depends on the second setting.
- Online Speech Recognition. Found at Settings → Privacy & Security → Speech. When this is on (the default), your audio is sent to Microsoft's cloud speech recognition service. When it's off, Voice Typing falls back to a local speech model built into Windows.
The distinction matters: most users who have never visited the Speech privacy settings are running Voice Typing with Online Speech Recognition enabled. Their dictation audio is going to Microsoft's servers on every use.
What Online Speech Recognition sends to Microsoft
With Online Speech Recognition enabled, Windows transmits:
- The audio stream of everything you dictate
- Contextual data Microsoft uses to improve recognition quality, including contact names and calendar entries from apps that integrate with Windows speech personalization
- A speech service ID associated with your Microsoft account (not an anonymous identifier — your voice data is linked to your account if you're signed in)
Microsoft's privacy statement states that speech data is used to improve Microsoft speech services and may be reviewed by employees or contractors for quality assurance. The retention window for voice data is not published as a specific number of days — Microsoft's documentation states data is retained "as needed to provide the service."
Unlike Apple's random identifier approach, online speech recognition in Windows is tied to your Microsoft account when you're signed in. This means dictated content is part of your Microsoft account activity history unless explicitly deleted via the Microsoft privacy dashboard.
The local speech recognition fallback
Turning off Online Speech Recognition at Settings → Privacy & Security → Speech switches Voice Typing to a local model. This keeps audio on your device. The tradeoffs:
- Accuracy. The local model is generally less accurate than the cloud model, particularly for specialized vocabulary, accents, and noisy environments.
- Language support. Local speech recognition is available only for a subset of the languages supported by the online service.
- No indicator. Like iOS, there is no per-session UI indicator showing whether a given dictation session used local or cloud processing. The setting is global, but third-party apps can also invoke speech recognition independently of this setting.
- Other apps may still use cloud. The Online Speech Recognition toggle controls Windows' built-in speech services, but applications that implement their own cloud speech (browser-based transcription, web apps, Office 365 dictation) operate independently of this setting.
Office 365 Dictation is a separate data path
Microsoft Word, Outlook, and other Office apps include their own Dictate feature, distinct from Windows Voice Typing. Office Dictate uses Azure Cognitive Services — a separate Microsoft cloud service with its own privacy policy and data retention terms.
The data handling for Office Dictate depends on your organization's Microsoft 365 licensing tier and whether your tenant has the Connected Experiences privacy controls enabled. Enterprise customers on Microsoft 365 E3/E5 with data residency add-ons get different routing than consumer Microsoft 365 Personal subscribers.
The practical implication: disabling Online Speech Recognition in Windows Settings does not affect Office Dictate. The two features are separate, controlled by separate settings, and have separate data handling policies.
What the documentation doesn't cover
Microsoft's privacy documentation has notable gaps for professional use cases:
- No published retention window for online speech data — "as needed to provide the service" is not an auditable commitment.
- No list of subprocessors for the online speech recognition pipeline. Azure's subprocessor list covers Azure services broadly, not Voice Typing specifically.
- Human review disclosure is present but not opt-outable. Microsoft states speech data "may be reviewed" for quality. There is no way to use online speech recognition while opting out of potential human review.
- HIPAA coverage for Voice Typing is not addressed. Microsoft's HIPAA coverage for Azure services doesn't extend to consumer speech features in Windows.
Comparison: Windows Voice Typing vs. zero-retention dictation
(For the same breakdown on iOS, see our iPhone dictation privacy analysis.)
| Feature | Windows Voice Typing (online) | Verity |
|---|---|---|
| Default audio routing | Microsoft cloud servers | Verity STT pipeline |
| Linked to user account | Yes (Microsoft account) | No — session-isolated |
| Published retention window | Not specified | Zero (deleted after transcription) |
| Contractual zero-retention | No | Yes — DPA with named processors |
| Human review possible | Yes (quality assurance) | No |
| Named subprocessors | Not published for Voice Typing | Published on privacy page |
| Used for model training | Yes (Microsoft speech improvement) | No |
What to do if you dictate sensitive content on Windows
- Turn off Online Speech Recognition at Settings → Privacy & Security → Speech if local accuracy is acceptable for your use case. This keeps Voice Typing audio on-device.
- Note that Office Dictate is separate. Disabling the Windows speech setting does not affect Office apps. Review your Microsoft 365 Connected Experiences settings separately if you use Office Dictate for sensitive documents.
- Use Verity for dictation that requires a privacy guarantee. Verity for Windows provides zero-retention dictation with a published, contractual data processing agreement. Audio is processed and immediately deleted — no retention, no account linkage, no model training. (Architecture details · Download Verity.)
Windows Voice Typing is a capable feature, and the local fallback option is a meaningful privacy improvement over the default. But "local fallback" is not the same as a privacy guarantee, and the cloud path — which remains the default — has significant gaps for anyone dictating professionally sensitive content.
Dictation that doesn't retain your audio
Zero-retention architecture · Named processors · Free to start
Download Verity