Skip to main content
DoneThat

AI Adoption GuideGeneralSelect

Dictation AIs

Dictation AIs turn speech into working text when speaking is faster than typing, especially for drafting, communication, accessibility needs, or getting past a blank page.

By Don, DoneThat’s AI coach · updated

When speaking is faster than typing

Use a dictation AI when talking gets words onto the page faster than typing. The output is a transcript you can replay against the audio. If there is no audio, the transcript stays empty. The tool does not invent a sentence the speaker did not say. You still edit. You still hit send.

That is the quality bar. Fluency is not. A paragraph that sounds finished can still be wrong.

Dictation earns its keep on a blank page, in a first draft of an email or note, in talking points you will tighten later, and when typing is slow, painful, or inaccessible. Speak if your hands are occupied and your thoughts are already in sentences. Type if you already know the exact wording and a missed "not" would change the meaning.

Voice also helps when you stall on the first sentence. Saying the messy version out loud often produces a paragraph you can cut down. That is still a draft. It is not a finished message, and it is not a quote.

It is the wrong tool when you need a clean quote for someone else, a legal record, or a meeting transcription of several people talking over each other. Those jobs need speakers labeled, a different capture setup, and a different review. Personal dictation is one voice and one working draft.

Treat the pass as input. The send button stays yours.

Record, then transcribe

Record first. Transcribe second. Keep the audio so you can replay it. A live caption you cannot replay is a convenience, not a source you can stand behind.

Use a room quiet enough that playback is audible. Background speech will show up as extra words, or as two voices you cannot separate. If someone else starts talking, stop the recording. Do not send that file through a personal dictation tool and call the result your words.

If the recording is silent, you get empty output. Leave it empty. Do not paste an old transcript, a guessed paragraph, or a cleaned-up version to fill the hole. Empty stays empty because there is nothing to check. Asking ChatGPT or Microsoft Copilot to "just write what I probably said" is inventing a sentence.

Keep sessions short enough that you will actually listen back. Name people and numbers slowly. Pause at sentence boundaries if your tool needs the gap. If you stop mid-thought, leave the halt in the audio. When you resume, start a new sentence. Restarting in the middle of a clause is how two intents get glued into one false statement.

If you are about to say a password, an access token, health information, a customer list, or anything your policy treats as restricted, stop talking. On-device options such as Mac dictation keep more of the work on the machine. Cloud tools in the same class, including Whisper, Wispr, Microsoft Copilot, and ChatGPT, may send audio or text off the device. If policy forbids that destination, do not dictate the secret into the tool.

Transcribe from the file you will replay. Do not merge a live partial with a second pass you typed from memory and then call the mix a transcript.

Edit the transcript against the audio

Replay the audio with the text on screen. Correct what was said. Do not upgrade it into what you meant to say, and do not write a sentence that is not in the recording.

Homophones, names, and dropped negatives are the usual misses. "Don't ship Friday" becoming "ship Friday" is a meaning change. If a clause is muffled and you cannot hear it, leave a gap or a timestamp. A confident invented clause is worse than an obvious hole.

Punctuation is part of the edit. A pause that was a comma can become a period that splits a qualifier off the claim. After you fix the words, read the sentences once. If a sentence would surprise the person who spoke, it is not done.

Strip filler that adds nothing. Keep hedging when the hedge is the point. Deleting "I think" or "we might" can invent certainty the speaker did not offer.

The same rule applies when you paste the transcript into AI assistants to shorten or restyle. Tone and length can change. A promise, a date, or a quote that was never spoken cannot appear. If the assistant adds one, you delete it.

Here is the one check that matters in practice. You walk between meetings and dictate a note for Priya: the launch copy is in review, you need her legal pass by Thursday, and you are not promising Friday. The transcript comes back tidy and drops "not." It now reads as if Friday is fine. You only catch it because you replay the clip. You restore the negation, cut the filler words, and then send. The usable artifact is the edited note. The raw dump is not.

If you later turn those talking points into slides, dictation can feed a spoken outline the same way vibe-coding slides starts from an outline you already trust. The quality bar does not change. Match the audio before the outline becomes a deck.

Send the edited text, not the raw dump

Do not treat raw dictation as sent. Speech is full of false starts, filler, and half sentences that listeners tolerate and readers do not. A dropped negative becomes a commitment. A name the model guessed becomes the wrong person on the thread. If you dictated while walking, re-hear names and numbers first, then the greeting and the ask, then send. Pasting into chat because the paragraph looks already written is how raw dictation leaves the building.

Do not treat the transcript as a quote either. A quote is wording you can defend as theirs, checked line by line. Dictation is a working draft. If you cannot replay, you do not have a quote. Inventing a quote from a fuzzy recording is the failure mode that looks professional and is still false.

After you send, keep or discard the audio under the same retention rules you use for other recordings. A tidy transcript is not permission to leave a cloud copy of a conversation your policy would have kept local.

Choosing a dictation tool as a class

Whisper, Wispr, Mac dictation, Microsoft Copilot, and ChatGPT all turn speech into text. They belong to one class. This page does not rank them.

What they share is the loop. You speak. You get words. You check those words against the audio. None of them is a send button. None of them is a witness. None of them should fill an empty recording with a plausible paragraph.

Where they differ as a class is mostly destination and how much they tidy as they go. On-device dictation is often the conservative default when policy is strict. Cloud assistants may punctuate, paragraph, and smooth. That help is useful for a first draft. It is dangerous if you confuse fluent prose with accuracy. A model that writes complete sentences can invent the sentence that was never there. Your check is the audio. If there is no audio, you have nothing to check, so you do not have a transcript to send.

Do not assume the assistant you already use for writing is the right place to put this audio. Microsoft Copilot and ChatGPT can sit in the same class as Wispr or a Whisper-based workflow and still be the wrong destination for the recording. Mac dictation may be enough for a short note you will edit in place.

Write the house rules once, in the place you already keep task instructions. An AI Projects note that says "never fill gaps, keep names, stop if this is restricted" beats switching tools every week.

Pick the tool your device, your policy, and your workflow already allow. Then run the same four steps: record, transcribe, edit against the audio, send.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first