Skip to content

What gapi is

gapi turns recordings of Uzbek conversations into text and tells you what happened in them: who spoke, about what, in what mood, and which personal data was said aloud.

How it works:

  1. You upload a recording, through the cabinet or the API.
  2. A few seconds later the result is ready: text by utterance with timing, speakers and the options you switched on.
  3. You read it in the cabinet, fetch it by link, or receive it through a webhook.

What it can do

CapabilityWhat you getTariff
Speech recognitionText with the time of every utterance. Any recording in Uzbek: a call, a meeting, an interview, a dictaphone; in Latin script×1
Term vocabularyCompany and product names and people's names are recognised more often×1
Punctuation, numbersText that reads as written: punctuation, capitals, digits×1
Who is speakingThe sides by stereo channel, or voices separated in mono×1
Voice gender, client and managerWhich party is which×1
Recognise by voiceThe same person gets the same number in every recording×1
EmotionsThe mood of every utterance and a summary per speaker×2
Personal dataNames, phones, addresses, documents found and masked×1.2

The tariff multiplies the length of the recording; it is not a charge per option. Every ×1 row can be on at once and an hour of audio is still an hour. Only Emotions and Personal data make it cost more.

Every option in detail: Options.

Where to start

  • Quick start: a token, the first recording, a result in five minutes.
  • Use cases: ready-made option sets for calls, meetings, interviews.
  • Result format: what the JSON holds.
  • API: every request, limits, webhooks.

What it costs

The unit is g: 1 g is a minute of audio at ×1. An hour of recording without paid options costs 60 g. Emotions double the bill, Personal data adds 20 %, and together they make ×2.2. During the beta every account's balance is topped up to 240 g, four hours of audio, every midnight. In detail: Balance and billing.