Skip to Content
Frequently asked questions
1. Pricing and FAQ

FAQ

This page answers frequently asked questions about mocoVoice.

Choose the topic closest to what you are looking for.

General

Can I record and transcribe with a microphone?

You can record audio using your web microphone or web browser. For detailed instructions, see the transcription guide pages.

Where do microphone recordings go?

Microphone recordings that have not been transcribed appear under “Select from Past Recordings” at the bottom of the file upload screen. You can select them from the list to start transcription.

How is mocoVoice different from other speech recognition services?

mocoVoice stands out in the following ways:

  • Simple API design for easy integration.
  • Transcribes video files as well as audio files.
  • Delivers high-accuracy transcripts.
  • Phonetic readings are optional when registering dictionary words.
  • Includes speaker separation features.
  • Removes filler words and produces proofread, high-quality transcripts.
  • Lets you set custom meeting notes templates.
  • No limits on meeting notes creation or translation requests, as long as server capacity permits.
  • Supports multiple languages, including Japanese and English, with automatic code-switching.

Will mocoVoice’s performance improve over time?

Yes. mocoVoice continuously trains on data and improves its architecture, delivering high-performance models that handle the latest terminology.

Does it recognize everyone’s voice?

Yes. It is designed to recognize voices from any speaker.

Does it handle different intonations and dialects?

Yes. You can also use dictionary registration and dedicated models to further improve accuracy.

Does it maintain accuracy if I speak word by word?

Yes. It transcribes single-word utterances without issues.

How does transcription work?

mocoVoice uses AI that processes both acoustic and linguistic information to transcribe audio. We also use large language models (LLMs) to further improve accuracy.

How secure is mocoVoice?

We store audio and transcript data in encrypted storage. Only authorized users can access it. You can also request an on-premises deployment.

Pricing and Contracts

Can I extend my mocoVoice trial?

Submit an extension request and your reason using the mocoVoice Support Form. A representative will contact you.

Do you offer custom pricing?

Yes, we offer custom contract plans. Request a custom contract through the mocoVoice Support Form.

How do you calculate fractions of a cent in billing?

We round up fractional amounts.

What happens after the Trial plan ends?

After your Trial plan ends, your team’s plan changes to “No plan registered”. Your past data remains, but you cannot use features. Subscribe to a paid plan to restore access to transcription and other features.

Why does it say “No plan registered”?

When a Trial plan ends or you create a new team, your plan changes to “No plan registered”. Your past data remains, but you cannot use features. Subscribe to a paid plan to restore access to transcription and other features.

I received a special offer at an exhibition. What should I do?

Enter the event name and your requested contract type in the “Other” field of the mocoVoice Plan Change Request Form. We will verify your name against the attendee list and contact you with a custom plan.

If I sign up mid-month, is the cost prorated?

Yes. If you sign up mid-month, your first month’s fee is prorated based on your start date.

How do I cancel my subscription?

Select “Request Cancellation” on the Plan Change Request Form and submit your request.

Technical Questions

Is the mocoVoice API as accurate as the mocoVoice app?

Yes. The mocoVoice API delivers the same high-accuracy transcription as the mocoVoice application.

How long does transcription take?

Depending on server traffic, the mocoVoice API can process a 1-hour audio file in about 5 minutes.

What should I do if the transcript accuracy is poor?

Recording environment and file format affect transcription accuracy. Check the following:

  • Recording environment: Ensure the audio is clear and free of noise.
  • File format: Use the recommended format (WAV, mono, 16kHz).
  • You can also improve accuracy by using dictionary registration or dedicated models.

What languages do you support?

We support about 90 languages, including Japanese (ja), English (en), Chinese (zh), German (de), Spanish (es), Russian (ru), Korean (ko), French (fr), Portuguese (pt), Turkish (tr), Polish (pl), Dutch (nl), Arabic (ar), Italian (it), Hindi (hi), Vietnamese (vi), and Thai (th). For a complete list of language codes, see the API Reference.

Do you support real-time transcription?

We are currently developing a real-time API. Stay tuned for future updates.

Can I send batch requests?

We are currently developing a batch request API. Stay tuned for future updates.

What is the API response format?

For details, see the mocoVoice API Reference.

Is there a file size limit for audio uploads?

You can upload files up to 5GB. However, depending on the number of registered words, files under 5GB might still fail to process.

Are there restrictions on data formats?

We support the following audio and video formats:

  • Audio: wav, mp3, m4a, caf, aiff, wma, flac, ogg, aac
  • Video: avi, mp4, rmvb, flv, mov, wm

We support sampling rates of 8kHz, 16kHz, 22.05kHz, 44.1kHz, 48kHz, and 96kHz in mono or stereo. The maximum file size is 5GB. For the best speed and accuracy, we recommend WAV format, mono, and 16kHz.

Does speaker separation reset after 3 hours?

mocoVoice processes speaker separation in 3-hour chunks. If your audio exceeds 3 hours, speaker IDs may change at the 3-hour mark.

For example, a speaker identified as “SPEAKER_00” in the first 3 hours might be labeled “SPEAKER_02” afterward, even though it is the same person. This is expected behavior. Check the transcript and use the speaker editing feature to assign the same name to the same person.

Can I transcribe video files?

Yes, you can transcribe video files.

What are the system requirements?

We recommend the following browsers:

  • Google Chrome (Windows / Mac)

Can I transcribe offline?

You need an internet connection to send transcription requests and view results. However, you can go offline while waiting for the results to process.

Is there a limit to how many words I can add to the dictionary?

The limit is 1,000 words. Exceeding 1,000 words does not cause an error, but it acts as a soft limit.

Do I need to add phonetic readings when registering words?

No, it is not required. However, adding phonetic readings improves recognition accuracy.

Can I register multiple phonetic readings for one word?

Yes. You can register multiple readings separated by a pipe (|), like “reading 1 | reading 2 | reading 3”. There is no limit to the number of phonetic readings.

How can I improve recognition accuracy when speaking?

Speak clearly and loudly in a quiet environment with no background noise to improve recognition accuracy.

Account and Other

I did not receive the account confirmation email.

Check your spam folder for emails from no-reply@mocomoco.ai. If you still cannot find it, contact us through the mocoVoice Support Form.

How do I reset my password?

Click “Forgot Password” on the mocoVoice login screen to reset your password.

How do I delete a team?

Submit a deletion request with your team name through the mocoVoice Support Form.

Where do you process my audio data?

We primarily use domestic and international servers. However, we offer custom plans that use only domestic servers. For details, contact us through the mocoVoice Support Form.

How can I check for service outages?

You can check the Service Status Page.

Do you offer an on-premises version of mocoVoice?

Yes. We offer custom contract plans for on-premises deployment. Contact us through the mocoVoice Support Form with your on-premises requirements.

Other Questions

If your question is not listed here, contact us through the mocoVoice Contact Form.

Last updated