Skip to content دوام Help Center

AI & Automation

Voice dictation

One-line summary: speak instead of typing — a button inside the composer turns your voice into text that appears in the field while you are still talking, and the audio itself is never stored and never leaves Duwam's own servers.

Plan: Available on Professional, Advanced and Business and during the trial (and on legacy Custom-plan subscriptions); not available on Basic, or once a trial ends without subscribing. It is also off by default even on the plans that include it: the platform operator switches it on for your company, and the transcription service has to be deployed — until both are true the button is not drawn at all. No permission is involved: whoever can type in the field can dictate into it. Start with Chat & messaging.

Overview

Dictation in Duwam is an input method, not a screen. There is no page to open and no menu to hunt through: it is a small button sitting among the composer's other buttons. You press it, you talk, and your words appear as text in the very field you were about to type in.

It is built on four promises, each of which cost something to keep:

  • The audio is never stored. It is processed in memory and discarded the moment the text comes back: never written to disk, never sent to any third party. Transcription runs on a service Duwam hosts itself.
  • The text is never re-written. What tidies it afterwards is a set of blunt rules that strip hesitation sounds and fix spacing — not an AI model rephrasing your message to your manager.
  • Nothing is written over what you typed. Insertion happens at the caret, and a space is added only where one is needed.
  • Nothing is sent. Dictation fills the field and stops; reading, editing and sending stay with you.

And before the microphone opens, one thing is required: your consent to how the audio is handled. That is not a browser dialog you can route around — it is a condition the server checks before it will accept any recording at all.

Where do I find the dictation button?

In three composers, and no fourth:

  1. Team Chat — the Human Resources workspace → the "Team Chat" row → the "Messages" tab. (For an employee it is a single row with no tabs that opens messages directly.) The button sits with the composer's other buttons.
  2. Inbox — the CRM workspace → the "Customer Chats" row → the "Inbox" tab, then the composer beneath the open conversation.
  3. The Duwam Assistant panel — at the foot of the panel, beside the question field.

The icon is a waveform, not a microphone, and the difference is deliberate: the microphone icon in the chat composer sends a voice note — the recording itself — while dictation turns speech into text. Two microphones in one composer meaning two different things is an interface that needs explaining.

The button is labelled "Start dictation" — that is what you read on hover and what a screen reader announces.

Note: The button is only drawn after the browser has confirmed both that the feature is open for your company and that the transcription service is live. So if you cannot find it, it was not drawn — you are not looking in the wrong place. A button that appears and then fails when pressed is worse than no button.

How do I dictate a message?

  1. Put the caret where you want the text to land. Insertion happens at the caret, so you can dictate into the middle of a sentence you typed by hand.
  2. Press "Start dictation" — or press Ctrl + Shift + M (⌘ + Shift + M on a Mac; the shortcut accepts either key). The first time, the browser asks for microphone permission.
  3. Talk. The button turns red and pulses. About two seconds into the recording the text starts arriving in batches without you stopping: whatever has settled is inserted into the field immediately, and whatever has not settled yet shows in a grey strip above the composer labelled "Dictation preview".
  4. Press the button again (or the same shortcut) to stop. Whatever speech is left is sent and inserted, and the microphone is released immediately.
  5. Read it, fix it, send it yourself.

Note: What lands in the field while you are talking is final and never re-written. Nothing settles until a second of later audio has passed over it and a real pause follows it — because a word still touching the edge of the recording is a word that has not been heard yet. That is why you never watch a sentence change under you after it appears. The grey strip is the only thing that moves, and it never enters the field.

Tip: Shorter recordings are both faster and more accurate. Live insertion does not even begin before the second second — so a very short clip arrives whole when you stop, which is quicker for it.

Caution: Esc cancels a recording in progress. Words that had settled and already appeared in the field stay — they were spoken and finalised — and only the unheard tail is discarded. A recording where nothing ever settled is discarded entirely.

And if no speech was picked up — silence, or noise — nothing at all is inserted and you read "No speech was detected.". That is deliberate: transcription models invent words over near-silent audio, and a field where the system writes words you never said is a message to your manager that you did not write.

The first time you dictate, this is the order of events: you press the button, you talk, you press stop — and then, instead of the text, a dialog opens titled "Before you start dictating". That first recording is discarded and never transcribed; the server refused it before reading it.

The dialog shows the consent notice exactly as the server sends it, one sentence per line. The dialog cannot shorten it or reword it, and if the text cannot be loaded there is no consent dialog at all — only one line: "The consent notice could not be loaded, and we cannot take your agreement without it. Please try again later." An "I agree" button over text nobody was shown is not consent, it is a button.

The notice says five things:

  1. When you use dictation, your voice is recorded in the browser and sent to a Duwam server to be turned into text.
  2. The audio is processed in memory and deleted as soon as the conversion finishes: it is not written to disk, not saved, and not sent to any external party.
  3. We do no speaker identification, we build no voice model of you, and we do not use your voice to train any model.
  4. Only the transcribed text is stored, in a log that belongs to you alone — your manager and your colleagues cannot see it — and you can delete any line of it, or all of it, at any time.
  5. Withdrawing consent stops dictation and erases your entire history immediately.

Beneath the text are two buttons — "I agree, start dictation" and "Not now" — together with the version tag under the label "Consent notice version". Once you agree the dialog closes and you read "Done. You can dictate now." — then press the button again and start, because your first recording went with the refusal.

Note: Consent is checked on the server, not in the browser. No recording is accepted from anywhere before your agreement is on record — not from an old tab, not from another window — and the refusal reads "You need to agree to how your audio is processed before dictating."

Caution: Your agreement is bound to the version of the text, not to the feature. If the notice changes materially — a new retention period, a new place the audio is processed — the old agreement no longer covers it and you are asked again; agreeing to text you were never shown is not agreement. And if it changes while you are reading it you are told "The consent notice changed. Please read it again, then agree." and the new text is shown.

The same dialog carries a third button for anyone who has already agreed: "Withdraw consent and erase my history". Pressing it does nothing on its own; it replaces the ordinary buttons with a plain question:

"Withdrawing consent stops dictation and erases every transcript saved in your history. This cannot be undone."

The only two answers on screen are then "Yes, withdraw and erase" and "Go back" — and neither sits where the first button was, so muscle memory cannot reach the destructive one. Afterwards you read "Consent withdrawn and your history erased.", and if something goes wrong, "Could not withdraw consent. Please try again." with the same question still in front of you.

Caution: Withdrawal is not just a stop for the future. It erases every line of your history permanently: no trash, no backup, no undo. That is on purpose — the notice promised it, and a promise in the copy that the code does not keep is worse than no promise.

Caution: That said, there is no route inside the product to reach that button today. The consent dialog only opens when the server refuses a recording for lack of a current agreement — and in that state you have nothing to withdraw, so the button is not offered; and once you agree, the dialog closes immediately. If you want to empty what has been stored for you, use "Delete all dictations" in the "Dictation" panel (below), and talk to your company's administrators about switching the feature off.

How do I choose the dictation language?

Beside the dictation button is a small chip reading AR or EN, labelled "Dictation language". Press it to flip between the two. Your choice is saved on this device, not on your account — every computer keeps its own, and the setting does not follow you to another machine.

  • Choose before you press. The language is sent with the recording; it is not guessed from it.
  • "Automatic" in the settings panel does not mean "detect the language from the audio". It means your interface language: Arabic if your interface is Arabic, English if it is English. Detecting from audio costs measurable time on every single request, and the chip is both faster and more honest.
  • The chip is visible from the first paint, not only while recording, so the choice comes before the speech rather than after it.

Note: Mixing the two languages inside one sentence is not guaranteed — the transcription model writes one language per passage. Product names in English, however, are restored to Latin script by a separate layer, whichever language you dictate in (see below).

What happens to the text before it reaches me?

Two layers run after transcription, and both are deterministic: no AI model in either, no cost to use them, and no ability to invent a word.

1. A product-name dictionary. It restores a spoken term to its known Latin spelling, and every company gets it with no setup: «واتساب» → WhatsApp, «قوقل» / «جوجل» / «غوغل» → Google, «اكسل» → Excel, «زوم» → Zoom, «تيمز» → Teams, «اوتلوك» → Outlook, «بوربوينت» → PowerPoint, «بي دي اف» → PDF, «شات جي بي تي» → ChatGPT. Matching is on whole words only, never a fragment, so an ordinary Arabic word is never rewritten because a term happens to sit inside it. The dictionary also understands the fused Arabic definite article, so «الرفتر» is read as "الـ report" and written that way.

2. A rule-based clean-up. It removes hesitation sounds only — «آآ», «ااا», «ممم», um, uh, erm, hmm — fixes spacing around punctuation, and normalises digits to Latin numerals.

Note: The clean-up is deliberately conservative. The Arabic word «يعني» is not removed, for instance, because it is a real word ("it means") that merely also gets used as filler — stripping it would change whole sentences. The rule is: when in doubt, leave it. Seeing an «آآ» and deleting it yourself is cheaper than losing a word you said without ever knowing.

Where do I read my dictation history?

Every dictation is stored as text only, in a log that belongs to you alone inside your company: your manager and your colleagues cannot see it, and there is no admin screen that reads it for anyone. It is less an archive than an undo — for the times the words landed in the wrong window, or were deleted a second too fast.

Open it from the Duwam Assistant panel → the small chevron button beside the dictation button, labelled "Dictation history and settings". A "Dictation" panel slides in with two sections: "History" and "Settings".

Caution: That is the only door to the history. The dictation button in Team Chat and in the Inbox does not carry that chevron — it was removed there on purpose because the composer is already crowded. So if Duwam Assistant is off for your company, there is no door to the history at all.

Each row shows the text, when it was said, the recording's length as minutes:seconds under the label "Recording length", and three buttons:

  • "Copy" — copies the raw transcript to the clipboard; you read "Copied" or "Couldn't copy — copy it manually".
  • "Insert into the field" — puts the stored words back into the composer. It does not re-record and does not transcribe anything again; the words already exist, and this simply hands them back.
  • "Delete" — removes the row.

Rows are ordered newest first, up to 200 of them. An empty history renders a sentence rather than a blank: "No dictations yet — what you dictate will appear here" — because a panel that paints nothing reads as a panel that failed to load, and the two states need opposite reactions. If loading really did fail you read "Couldn't load the dictation history."

To clear the whole log: the "Delete all dictations" button under the list — offered only when there is something to clear — which asks "Delete all dictations? This cannot be undone." with "Yes, delete all" or "Cancel".

Tip: Deleting one row happens in front of you immediately. If the server refuses it, the row comes back to the same seat — not to the end of the list — and you read "Couldn't delete — the row is back." Nothing that disappears from the screen survives quietly behind it.

What can I change in "Settings"?

Three fields in the lower half of the "Dictation" panel, all saved on this device rather than on your account — because a chosen microphone is a property of the machine, not of the person:

  • "Dictation language" — "Automatic", "Arabic" or "English". The same choice the chip beside the button flips.
  • "Text clean-up" — "On" or "Off".
  • "Keep history for (days)" — the default is 30, and the highest the field accepts is 365.

Under the last field is a line that hides nothing: "You can delete your history at any time. Automatic deletion after this period is not active yet." The number is a preference, not a guarantee today: nothing sweeps your history by age, and deleting is your own doing.

Caution: "Text clean-up" has no effect today either. None of the three composers sends that choice with the recording, so clean-up runs always, even when you have selected "Off".

What if the microphone does not work?

Each cause has its own message, because "something went wrong" is not a message but a shrug — and it cannot tell "you denied this" from "another app is holding the device":

What you read What it means and what to do
Microphone access was blocked. Open your browser's permission settings, allow the microphone for this site, and try again. The permission is denied in the browser — open it from the address bar and allow the site.
No microphone was found. Connect one, or choose another device in the dictation settings. No recording device is attached.
The microphone is busy in another program. Close the app using it and try again. A call or recording app is holding the device.
The selected device is unavailable. Choose another microphone in the dictation settings. The requested device is no longer there.
Microphone access is disabled on this page. Ask your system administrator. A device- or network-level policy is blocking recording.
Dictation needs a secure connection (HTTPS). Open the site over a secure link and try again. The page is open over an insecure connection.
The microphone could not be started. Check your permissions and try again. An unclassified cause — start with permissions.
"Could not transcribe the audio. Please try again." The recording itself was not transcribed: too long for the limit, or the service did not answer.

Note: The seven microphone messages above are served in Arabic today, in both interface languages; the English is a translation so you can recognise them. The last row is the only one of the eight that is translated.

Note: Two of those messages point at choosing a device in "the dictation settings" — a device picker that does not exist in the settings panel today. Change the default device from your operating system or your browser settings instead.

Two limits explain most "Could not transcribe the audio" cases: a single recording may not exceed 10 MB, and each transcription call has a 30-second timeout. Several short clips are safer than one long one.

What dictation does not do

So you do not go looking for what is not there:

  • It does not store your voice or lend it to anyone. The audio is processed and dropped; there is no copy in any log and none with any third party.
  • No speaker identification, no voiceprint, and no training of any model on your voice.
  • No supervisor screen. There is nowhere in Duwam to read what your team dictated — not for a manager, not for the account owner. The history belongs to its owner alone.
  • It never sends on your behalf, never presses send, and never creates a task or a request. It fills the field and stops.
  • It never touches the clipboard. Whatever you had copied before dictating is still there — unless you pressed "Copy" in the history yourself.
  • It is not in every field in Duwam. Three composers only: no dictation in a task title or description, in form fields, or in search.
  • No voice commands. Nothing you say means "new line" or "delete that"; what you say is written as you said it.
  • No microphone picker and no editable shortcut inside the product; the shortcut is fixed at Ctrl/⌘ + Shift + M.
  • No company word list. The product-name dictionary ships as it is, and there is no screen today for adding your own words or your employees' names to it.
  • No automatic deletion by age, no archiving, and no export of the history.
  • It does not work over an insecure connection, nor in a browser that cannot record audio.
  • It does not spend your monthly AI message allowance — transcription never passes through a model billed per message. See Duwam AI Assistant.

FAQ

Q: I talked, pressed stop, and got a consent notice instead of my text — where did my recording go? A: It was refused before it was read. Consent is a condition the server checks, not a browser dialog, so that first recording is gone. Agree, then press the button again and repeat the sentence.

Q: Can my manager hear what I dictate? A: No. The audio is never stored at all, and the text goes into a log only its owner can read — there is no admin screen that shows it.

Q: Is my voice kept anywhere after transcription? A: No. It is processed in memory and dropped the moment the text comes back: no disk, no backup, no third party.

Q: I dictated in Arabic and an English word came out in Arabic letters. A: That is how transcription models behave — one language per passage. That is exactly why the text then passes through the product-name dictionary, which restores the known ones to Latin script. A word that is not a product name needs a fix by hand, and there is no screen today for adding your own words to the dictionary.

Q: Why is there no dictation button on tasks or notes? A: Because there is none. Dictation lives in three composers only: Team Chat, the Inbox, and the Duwam Assistant panel.

Q: I do not see the dictation button at all. A: One of three things: your plan does not include it (Basic, or a trial that ended without subscribing), the platform operator has not switched it on for your company — it is off by default — or the transcription service is not deployed. The button is not drawn until both conditions hold.

Q: The dictation button works in Team Chat but I cannot find the chevron that opens the history. A: The chevron exists only in the Duwam Assistant panel. It was deliberately removed from the Team Chat and Inbox composers.

Q: I chose "Off" for "Text clean-up" and the text is still being cleaned. A: Correct. The choice is stored on your device but is not sent with the recording today, so clean-up always runs. In any case it only strips hesitation sounds and cannot change a word you said.

Q: I set "Keep history for" to 7 and nothing was deleted after a week. A: The field is a preference, not a guarantee — the line beneath it says so outright. There is no deletion by age yet; delete from "History" yourself, row by row or with "Delete all dictations".

Q: How do I withdraw consent and erase everything in one go? A: The button exists inside the consent dialog, but no route in the product opens that dialog for you once you have agreed. What is available today is "Delete all dictations" in the "Dictation" panel, plus a conversation with your company's administrators about switching the feature off.

Q: I pressed Esc mid-sentence and some of the text stayed. A: That is intended. What had settled and appeared in the field was spoken and finalised, and is not erased; only the unheard tail is discarded.

Q: Does dictation work on a phone? A: It works wherever the composer itself works and the browser allows the microphone, and only over a secure (HTTPS) connection.

Quick reference

Action Path
Start and stop dictation The "Start dictation" button in the composer — or Ctrl/⌘ + Shift + M
Cancel a recording in progress The Esc key
Switch dictation language The AR / EN chip beside the button
Dictate in Team Chat Human Resources → "Team Chat" → "Messages" → the composer
Dictate in the Inbox CRM → "Customer Chats" → "Inbox" → the composer
Dictate in the assistant The "Duwam Assistant" panel → the question field
Open the history and settings The "Duwam Assistant" panel → the chevron beside the dictation button ("Dictation history and settings")
Re-insert an earlier transcript The "Dictation" panel → "History" → "Insert into the field"
Copy an earlier transcript The "Dictation" panel → "History" → "Copy"
Delete one row The "Dictation" panel → "History" → "Delete"
Delete the whole history The "Dictation" panel → "Delete all dictations" → "Yes, delete all"
Set language, clean-up and retention The "Dictation" panel → "Settings"
Read the consent notice The "Before you start dictating" dialog — shown at your first dictation, before you agree