An abstract image in which a waveform-like line passes through several geometric gates and splits into three paths leading toward an arrow, a grid, and a connected structure.
AI and LLMsthrough Where the Work Breaks

Multilingual Voice Notes Can Now Trigger Field Follow-up Work

As multilingual speech recognition becomes available, small services can turn what people say in the field directly into work orders, handover notes, and confirmation requests.

Published 2026. 8. 27.

The foundation for speech to become work data

On August 26, 2026, Google introduced Gemini 3.5 Transcribe. It is still in public preview, meaning its access terms and features may change. It is not a finished consumer recording service, but a speech-recognition capability that developers can connect to their own services.

According to Google’s model documentation, it recognizes more than 85 languages, including Korean, as well as regional expressions. A language can be specified in advance, and it is designed to handle switches during a conversation, such as speech that mixes Korean and English.

In live mode, it displays captions as people speak, and Google states latency of under one second. However, each connection can last up to 10 minutes, and live mode does not provide speaker diarization—the ability to separate who said what.

It can process recorded files of up to one hour. That limit falls to 30 minutes when speaker diarization or word-level timestamps are enabled. The documentation says it can separate up to eight speakers, but separation for more than three speakers is experimental, so meetings where several people speak over one another need separate testing.

For words it often gets wrong, such as company names, part names, and field terminology, users can add up to 1,000 terms in advance. Google recommends typically using fewer than 100, so it makes more sense to start with the expressions essential to one industry than to add every possible term.

Google’s pricing page estimates paid use at about US$0.009 per minute for audio input and text output combined. That works out to about US$0.54 per hour, but this is an estimate based only on audio input and text output, and pricing may change during public preview.

The next screen matters more than the transcript

Consider an owner of a 12-person heating, ventilation, and air-conditioning installation and repair business. Three of the staff are more comfortable speaking Vietnamese than Korean. Every day, the owner deals with customer calls, site photos, technicians’ voice reports, and quotes spread across different places.

Customer requests are written down during calls on paper or in a phone notes app. The owner sends the details to technicians through messaging, technicians send photos and voice messages back from the site, and office staff listen to them and copy the information into the customer-management screen and quotes.

The same address and model name are entered several times. Letting technicians explain work in the language most comfortable for them speeds up communication, but office staff must interpret it again. If part quantities or visit arrangements are unclear, someone must call the customer once more.

A new service would not simply save customer calls or field reports as text. It would identify the customer request, product model, required parts, agreed date, owner, and items that need confirmation, then turn them into a single work card.

For live calls, captions should be used only as supporting material. After the call ends, the recording can be processed again to separate speakers. A separate translation function can preserve both the original and translated text, making it possible to check where meaning changed.

Technicians could also speak after finishing a job: “Checked outdoor-unit noise, replaced the drainpipe, return next week for the filter.” The service would separate completed work, additional work, and items requiring customer confirmation. It would ask the technician again only about high-risk details such as amounts, quantities, and dates.

Office staff would not read long transcripts from the beginning. They would review and correct only cards where work has stopped, such as “part quantity unconfirmed,” “customer visit-date approval pending,” or “no owner assigned.” Only confirmed information would move into a quote or work schedule.

Some work must clearly remain with people. Consent procedures for recording, confirmation of amounts and specifications, identifying speakers where several people spoke over one another, and final judgments involving safety and responsibility should not be automated away.

The completion condition for this service is therefore not “a transcript was created.” A case is complete only when an owner and deadline are set, important numbers have been checked by a person, and the information that the customer or next worker needs has actually been delivered.

Outside Korea, voice is used to connect work rather than merely document it

SUEZ (수에즈), a French water and waste management company, added Vivoka’s voice-input capability to an Android work app used by field technicians. When technicians spoke on site about meter replacements, leaks, reconnections, equipment numbers, and measurements, the information went into existing report fields.

Vivoka described the system as intended for more than 1,000 technicians and projected savings of 10 to 45 minutes per person per day. This is a supplier projection, not a verified result. Its multilingual recognition supports 41 languages, but it has not been disclosed whether SUEZ used multiple languages.

West Park Care (웨스트 파크 케어) in the United Kingdom worked with Spicerack (스파이스랙) on an app that turns voice notes from more than 100 visiting care workers into care records and provides translation for staff who are less comfortable with English. The work began by showing schedules and past care records, then passing results into an existing care-management tool.

Spicerack reported that, in a three-month trial, visit preparation time fell from eight minutes to two minutes and record accuracy rose from five out of 10 to seven out of 10. This is a supplier case study and needs independent verification. Still, it clearly shows that the connection to existing record screens—not speech recognition alone—produced the result.

Four small things to build now

1. A field-repair work card

  • This service separates customer calls and technicians’ field voice notes into work instructions, part requests, and return-visit schedules.
  • It is for repair businesses with five to 20 staff that work with foreign technicians and communicate schedules through calls, messaging, and paper.
  • It can accept Korean and multiple other languages in one function, while company and part names can be registered in advance, making it easier to test in one industry first.
  • The first screen should show customer cards for today’s visits, a large record button, and amounts, quantities, or dates that are still unconfirmed.

2. A visiting-care handover inbox

  • This service separates a care worker’s post-visit voice note into meals, medication, pain, unusual observations, and information for the next visitor.
  • It is for small visiting-care organizations that employ staff of multiple nationalities and pass a client’s condition through messaging during shift changes.
  • Records spoken in a worker’s preferred language can retain both the original and translated text, but health and medication information should be shared only after approval by the responsible person.
  • The first screen should show each client’s most recent visit record and important unresolved items, such as “medication not confirmed” and “not yet read by next worker.”

3. A safety-instruction repeat-back record

  • This service shows a manager’s safety instructions in a worker’s language, asks the worker to say back what they understood, and checks the difference between the two statements.
  • It is for small factories or construction businesses that explain daily work locations and hazards to foreign workers.
  • Multilingual transcription alone cannot determine whether someone understood. The narrow feature needed is a record of the translated instruction and the worker’s repeat-back.
  • The first screen should show today’s work team, selected language, a safety-instruction recording button, and whether each worker has completed the repeat-back.

4. A customer-interview commitment tracker

  • This service separates pain points, direct quotes, requested features, and follow-up dates from customer interviews, turning them into follow-up work.
  • It is for early-stage companies of around 10 people, and research agencies, where leaders meet customers directly but do not have time to replay recordings.
  • Recordings of conversations between two people place a relatively smaller burden on speaker separation. Registering product names and industry terms in advance can produce a result different from a general meeting-summary tool.
  • The first screen should show each interview’s core problem, statements whose evidence has not yet been checked, and follow-up commitments without an owner.

Today, check only the moments when people copy information

Ask the owner of one heating, ventilation, air-conditioning, or equipment-repair business to show, on a phone screen, how three recent jobs moved from the customer call to technician instructions and the quote. If, in at least two of the three cases, someone replayed audio or copied the same information into another screen, and the process required even one confirmation call, a field-repair work card is worth testing.

Why this matters where you are

Check whether field teams in your market replay voice messages, copy the same details into several systems, or make confirmation calls because key details are unclear. The language mix, recording-consent process, and systems that receive the final records may differ from the examples here. Start by testing one workflow where a spoken report must reliably become an assigned task, a date, and a human-checked detail.

Sources

6 sources

Every fact in this article came from the pages below. Check them yourself.

Multilingual Voice Notes Can Now Trigger Field Follow-up Work | Prometheon