voixa

package module
v0.2.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 3, 2026 License: MIT Imports: 16 Imported by: 0

README

voixa-go

Go client for the Voixa AI speech API. Standard library only; Go 1.22 or newer.

Text to speech in ten languages — built-in voices, voice cloning from a ten-second sample, voice design from a written description, podcasts up to 100 000 characters, streaming Vietnamese speech — and speech to text: audio or video in, transcript with timestamps, SRT and VTT out.

go get github.com/vovix-ai/voixa-go

Quick start

package main

import (
	"context"
	"fmt"
	"log"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("") // reads VOIXA_API_KEY; create a key in Studio → API keys

	voices, err := vx.Voices(ctx, voixa.English)
	if err != nil {
		log.Fatal(err)
	}
	clip, err := vx.Say(ctx, voixa.SpeakRequest{
		VoiceID: voices[0].VoiceID,
		Text:    "The lighthouse keeper switched the lamp on at dusk.",
		Project: "episode-12", // optional: group clips by content
	})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(clip.URL, clip.DurationSec) // signed WAV URL, valid for 1 hour
}

Say returns when the clip is ready. Use Speak with Wait: voixa.Bool(false) plus WaitForClip to manage the wait yourself — the API answers in under a second and you poll.

Every method takes a context.Context first; cancelling it stops a request or a wait immediately.

Client options
vx := voixa.NewClient(apiKey,
	voixa.WithBaseURL("https://api.voixa.vovix.io/v1"), // the default
	voixa.WithTimeout(60*time.Second),                   // per API call; default 30 s
	voixa.WithHTTPClient(myHTTPClient),                  // proxies, tracing, pacing
	voixa.WithUserAgent("my-app/1.2"),
)

An empty key falls back to the VOIXA_API_KEY environment variable. A client without credentials returns an error on every call instead of panicking.

Streaming (Vietnamese)

Hear the first words in about two seconds: audio arrives clause by clause while the rest is voiced. The stream is an io.ReadCloser; close it when done.

s, err := vx.Stream(ctx, voixa.StreamRequest{VoiceID: "vi-truc-ly", Text: "Xin chào, Voixa có thể giúp gì?", SampleRate: 16000})
if err != nil {
	log.Fatal(err)
}
defer s.Close()
io.Copy(player, s) // 16-bit little-endian mono PCM at s.SampleRate

SampleRate: 48000, 24000 (default), 16000 or 8000; Format: voixa.StreamPCM (default) or voixa.StreamWAV. Costs 1.5 tokens per character, charged when the stream opens (s.Tokens). The stream lasts as long as the context you pass; the per-call timeout does not cut it off.

Tokens

Every request is priced in tokens: 1 token = 1 Vietnamese character. Free tokens refill daily; top-ups carry over.

acc, _ := vx.Account(ctx)
w := acc.Tokens // Available, FreeToday, DailyFree, Balance, UsedToday, RefillsAt, Rates
w.Rates.EstimateSpeak(voixa.Vietnamese, 1200) // → 1200
w.Rates.EstimateStream(500)
w.Rates.EstimateTranscription(3600) // seconds of audio
w.Rates.EstimateClone()
w.Rates.EstimateDesign(3)

page, _ := vx.Usage(ctx, voixa.PageOptions{Limit: 20}) // every spend, refund and top-up

A request you cannot afford fails with an *voixa.Error of status 402 and code INSUFFICIENT_TOKENS (voixa.IsInsufficientTokens(err)); Needed, Available and RefillsAt tell you by how much and until when. Failed clips and voices are refunded.

Voices and samples

langs, _ := vx.Languages(ctx)                  // codes, labels, limits, clone-recording guides
samples, _ := vx.Samples(ctx, voixa.French)    // built-in voices with public sample URLs
voices, _ := vx.Voices(ctx, "")                // built-in voices, your clones, shared clones
voice, _ := vx.GetVoice(ctx, "vi-truc-ly")

Voice design: a voice from a description

No recording needed. Describe the voice, listen to up to three options, keep the one you like. A saved design is an ordinary voice of your account: use it with Say, MakePodcast and Stream. Characters for dubbing, narrators, ads — one voice per role.

d, err := vx.DesignVoice(ctx, voixa.DesignVoiceRequest{
	Description: "An old man in his seventies, hoarse and deep, speaking slowly",
	Language:    voixa.Vietnamese, // any language Languages() marks Designable
	Count:       3,                // 1–3 options, Rates.Design tokens each
})
done, err := vx.WaitForDesign(ctx, d.DesignID) // about a minute per option
for _, c := range done.Candidates {
	fmt.Println(c.Index, c.Status, c.URL) // listen, then pick
}

voice, err := vx.SaveDesign(ctx, done.DesignID, voixa.SaveDesignRequest{Candidate: 1, Name: "Grandpa Ba", Gender: voixa.GenderMale})
_, err = vx.Say(ctx, voixa.SpeakRequest{VoiceID: voice.VoiceID, Text: "Ngày ấy, ông vẫn nhớ con đường làng."})

// Or in one call, keeping the first option:
narrator, err := vx.CreateVoiceFromDescription(ctx, voixa.VoiceFromDescriptionRequest{
	Description: "A calm female narrator, warm and clear", Language: voixa.English, Name: "Narrator",
})

Options are drafts kept for 23 hours (ExpiresAt). Saving one costs Rates.Clone tokens. GetDesign, ListDesigns and DeleteDesign manage your designs; deleting a design keeps the voices saved from it.

Voice cloning and sharing

f, _ := os.Open("reference.wav") // 3–10 s of clean speech, WAV or MP3, max 10 MB
defer f.Close()
voice, err := vx.CreateVoice(ctx, voixa.CreateVoiceRequest{
	Name:     "Studio narrator",
	Language: voixa.English,
	Gender:   voixa.GenderFemale,
	Audio:    f, // any io.Reader; Size is optional for files and in-memory readers
})
_, err = vx.Say(ctx, voixa.SpeakRequest{VoiceID: voice.VoiceID, Text: "Hello from my own voice."})

err = vx.UpdateVoice(ctx, voice.VoiceID, voixa.VoiceUpdate{Shared: voixa.Bool(true)}) // let every Voixa account use it
err = vx.DeleteVoice(ctx, voice.VoiceID)

CreateVoice uploads the recording, registers the voice and waits for its sample. Voice cloning is available for Vietnamese, English, French, German, Italian, Spanish and Portuguese. Chinese, Japanese and Korean use built-in voices only; Languages reports Cloneable per language and CreateVoice rejects those languages with a 400.

Podcasts

Long scripts, up to 100 000 characters, become one episode: Voixa splits the text into sentences, produces the parts in parallel and joins them into a single WAV. Blank lines become short pauses.

episode, err := vx.MakePodcast(ctx, voixa.PodcastRequest{VoiceID: "vi-truc-ly", Title: "Episode 12", Text: script, Project: "my-show"})
episode.URL          // one WAV, signed for an hour
episode.DurationSec  // e.g. 4800 for an 80-minute episode

// or without waiting:
clip, err := vx.CreatePodcast(ctx, voixa.PodcastRequest{VoiceID: "en-alba", Text: script})
clip.Parts, clip.PartsDone // progress while Status is processing

MakePodcast waits up to two hours by default; check Status on the result.

Speech to text

Audio (or video — only the audio track is read) becomes a transcript with timed segments, SRT and VTT.

f, _ := os.Open("interview.mp3")
defer f.Close()
t, err := vx.TranscribeFile(ctx, voixa.TranscribeRequest{
	Audio:    f,
	FileName: "interview.mp3",
	Title:    "Interview with the lighthouse keeper",
	// Language: "vi",              // optional — Voixa detects it well on its own
	// Task: voixa.TaskTranslate,   // recognise any language, return English
	// Words: true,                 // per-word timestamps
	// Prompt: "Voixa, Vovix",      // spelling hints for names and jargon
})
fmt.Println(t.Language, t.DurationSec, t.Text)
fmt.Println(t.Segments[0])  // {ID Start End Text}
fmt.Println(t.SRTURL)       // signed SRT/VTT/TXT/JSON URLs, valid for 1 hour

Transcription runs on a machine that starts on demand, so it is never synchronous: Transcribe returns StatusProcessing as soon as the upload is done and you poll, while TranscribeFile does the polling for you (two to three minutes before recognition starts on a cold queue, then roughly one minute of work per five minutes of audio).

t, err := vx.Transcribe(ctx, voixa.TranscribeRequest{Audio: f, FileName: "talk.m4a"})
done, err := vx.WaitForTranscript(ctx, t.TranscriptID)
// progress while processing: done.DoneSec / done.DurationSec, done.JobStatus

page, err := vx.ListTranscripts(ctx, voixa.ListTranscriptsOptions{Project: "episode-12"}) // no text in lists
full, err := vx.GetTranscript(ctx, page.Items[0].TranscriptID)                             // text + segments + URLs
n, err := vx.DeleteTranscripts(ctx, []string{t.TranscriptID})                              // also deletes the upload

The container format is inferred from FileName (or the name of an *os.File); set Format when the name does not say. A reader whose length cannot be known (not a file, bytes.Reader, bytes.Buffer or strings.Reader, and no Size) is read into memory before upload. Languages you may pin, the accepted formats and the limits come from STTLanguages. Transcription is charged per second of audio when it finishes.

Library

projects, _ := vx.Projects(ctx)                                                       // groups with clip counts
page, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Project: "episode-12"})           // newest first
fromAPI, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Source: voixa.SourceAPI})      // made with an API key (vs. Studio)
podcasts, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Kind: voixa.ClipKindPodcast})
next, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Project: "episode-12", Cursor: page.Cursor})

shareURL, _ := vx.ShareClip(ctx, clipID) // public page, no account needed
_ = vx.UnshareClip(ctx, clipID)          // the link stops working
n, _ := vx.DeleteClips(ctx, []string{clipID})

Errors and limits

Every API failure is a *voixa.Error with Status, Message, Code (NOT_APPROVED, INSUFFICIENT_TOKENS, CLONE_LIMIT, VOICE_NOT_READY, INVALID, AUDIO_TOO_LARGE, NO_AUDIO; QUOTA and AUDIO_QUOTA are legacy) and RetryAfter. The client adds TIMEOUT (a wait gave up while the work was still processing — the work continues), SYNTH_FAILED, TRANSCRIBE_FAILED and DESIGN_FAILED.

_, err := vx.Say(ctx, req)
var ve *voixa.Error
switch {
case voixa.IsInsufficientTokens(err):
	// top up, or wait for ve.RefillsAt
case voixa.IsRateLimited(err):
	// wait ve.RetryAfter
case errors.As(err, &ve):
	log.Printf("%d %s: %s", ve.Status, ve.Code, ve.Message)
case errors.Is(err, context.Canceled):
	// you cancelled
case err != nil:
	// network error
}

Helpers: IsInsufficientTokens, IsNotFound, IsRateLimited, IsTimeout, ErrorCode, ErrorStatus.

  • One Speak call takes up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Whitespace and line breaks collapse to one space before counting; the client checks the length before sending (CollapseSpace, TextLength).
  • One podcast takes up to 100 000 characters (NormalizePodcastText).
  • One transcription takes up to 5 GB and eight hours of audio.
  • API keys are rate-limited to 1 request/second and 1 000 calls/day; new keys activate after two to three minutes.
  • Speech, streaming, transcription, designs and clones are paid in tokens (see Tokens above); read the wallet with Account.
Default waits
Method Poll interval Gives up after
WaitForClip, Say, WaitForVoice, CreateVoice, SaveDesign 2 s 3 minutes
WaitForDesign 4 s 10 minutes
MakePodcast, WaitForTranscript, TranscribeFile 5 s 2 hours

Pass voixa.WaitOptions{PollInterval: …, Timeout: …} as the last argument to change them; zero fields keep the default.

End-to-end check

cmd/e2e walks every public method against a live API with your key: account and usage, languages, samples, voices, Speak/Say with a project, clip listing and paging, projects, share links, cloning from a built-in sample, renaming and sharing the clone, a podcast, streaming, voice design, transcribing a generated clip back to text, error paths, then it deletes what it created.

VOIXA_API_KEY=… go run ./cmd/e2e
# another deployment:
VOIXA_API_URL=https://…/v1 VOIXA_API_KEY=… go run ./cmd/e2e

The speech-to-text step waits for a machine to start — five to eight minutes on a cold queue; VOIXA_E2E_SKIP_STT=1 leaves it out. VOIXA_E2E_SKIP_DESIGN=1 and VOIXA_E2E_SKIP_STREAM=1 leave out voice design and streaming. Exit code 1 if any step fails; cleanup still runs.

Documentation

Overview

Package voixa is the official Go client for the Voixa AI speech API.

Text to speech in ten languages (built-in voices, voice cloning from a short sample, voice design from a written description, podcasts up to 100 000 characters, streaming Vietnamese speech) and speech to text: audio or video in, transcript with timestamps, SRT and VTT out.

The package has no dependencies outside the standard library.

Authentication

Machine callers use an API key (create one in Studio, API keys). The key is sent in the x-api-key header. NewClient falls back to the VOIXA_API_KEY environment variable when the key argument is empty. A signed-in web user can be represented with WithIDToken instead.

Two kinds of work

Text to speech (Client.Speak, Client.CreatePodcast) answers quickly on a warm engine, so the API waits a little for you. Speech to text (Client.Transcribe) is as long as the audio, so it never waits: you get StatusProcessing and poll, or let Client.TranscribeFile do it.

How Speak behaves

The API synthesises in the background and waits up to about 20 seconds for the clip. On a warm engine that is a few seconds and the clip comes back ready with a URL. On a cold engine it may come back processing; Client.Say handles that by polling, Client.Speak leaves it to you.

Example
package main

import (
	"context"
	"fmt"
	"log"
	"os"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient(os.Getenv("VOIXA_API_KEY")) // create a key in Studio, API keys

	voices, err := vx.Voices(ctx, voixa.English)
	if err != nil {
		log.Fatal(err)
	}
	clip, err := vx.Say(ctx, voixa.SpeakRequest{
		VoiceID: voices[0].VoiceID,
		Text:    "The lighthouse keeper switched the lamp on at dusk.",
		Project: "episode-12", // optional: group clips by content
	})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(clip.URL, clip.DurationSec) // signed WAV URL, valid for 1 hour
}

Index

Examples

Constants

View Source
const (
	// MaxSpeakChars is the hard limit for one Speak call. Chinese, Japanese and
	// Korean have a lower limit per call (see Languages).
	MaxSpeakChars = 5_000
	// MaxPodcastChars is the character limit of one podcast.
	MaxPodcastChars = 100_000
	// MaxAudioBytes is the largest file Transcribe accepts: the most a single
	// upload can carry.
	MaxAudioBytes int64 = 5 * 1024 * 1024 * 1024
	// MaxAudioMinutes is the longest audio one transcription accepts.
	MaxAudioMinutes = 480
)

Limits enforced by the API, checked locally before a request is sent.

View Source
const (
	// CodeNotApproved: the account has not been approved yet (403).
	CodeNotApproved = "NOT_APPROVED"
	// CodeInsufficientTokens: not enough tokens for this request (402). See
	// Error.Needed, Error.Available and Error.RefillsAt.
	CodeInsufficientTokens = "INSUFFICIENT_TOKENS"
	// CodeQuota: legacy daily character limit reached (429).
	CodeQuota = "QUOTA"
	// CodeCloneLimit: voice clone limit reached (403).
	CodeCloneLimit = "CLONE_LIMIT"
	// CodeVoiceNotReady: the voice is still processing or failed (409).
	CodeVoiceNotReady = "VOICE_NOT_READY"
	// CodeInvalid: request validation failed (400).
	CodeInvalid = "INVALID"
	// CodeAudioQuota: legacy daily transcription limit reached (429).
	CodeAudioQuota = "AUDIO_QUOTA"
	// CodeAudioTooLarge: the audio file is larger than one upload can carry (400).
	CodeAudioTooLarge = "AUDIO_TOO_LARGE"
	// CodeNoAudio: nothing was uploaded for this transcript (400).
	CodeNoAudio = "NO_AUDIO"

	// CodeTimeout: a Wait helper gave up while the work was still processing (504).
	CodeTimeout = "TIMEOUT"
	// CodeSynthFailed: Say finished with a failed clip (502).
	CodeSynthFailed = "SYNTH_FAILED"
	// CodeTranscribeFailed: TranscribeFile finished with a failed transcript (502).
	CodeTranscribeFailed = "TRANSCRIBE_FAILED"
	// CodeDesignFailed: CreateVoiceFromDescription finished with a failed design (502).
	CodeDesignFailed = "DESIGN_FAILED"
)

Machine-readable error codes that accompany the message in Error.Code.

View Source
const DefaultBaseURL = "https://api.voixa.vovix.io/v1"

DefaultBaseURL is the public Voixa API.

View Source
const DefaultTimeout = 30 * time.Second

DefaultTimeout is the per-request timeout for API calls. It does not apply to audio uploads or to the body of a speech stream, which last as long as they need.

View Source
const Version = "0.2.0"

Version is the version of this client, sent in the User-Agent header.

Variables

This section is empty.

Functions

func Bool

func Bool(b bool) *bool

Bool returns a pointer to b, for optional boolean fields such as SpeakRequest.Wait.

func CollapseSpace

func CollapseSpace(s string) string

CollapseSpace applies the rule Speak and Stream use before counting: every run of whitespace, line breaks included, becomes one space, and the ends are trimmed.

func ErrorCode

func ErrorCode(err error) string

ErrorCode returns the Code of a *Error anywhere in err's chain, or "".

func ErrorStatus

func ErrorStatus(err error) int

ErrorStatus returns the HTTP status of a *Error anywhere in err's chain, or 0.

func IsInsufficientTokens

func IsInsufficientTokens(err error) bool

IsInsufficientTokens reports whether err means the account does not have enough tokens for the request (HTTP 402).

Example
package main

import (
	"context"
	"errors"
	"fmt"
	"log"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	_, err := vx.Say(ctx, voixa.SpeakRequest{VoiceID: "vi-truc-ly", Text: "Xin chào!"})
	var ve *voixa.Error
	switch {
	case voixa.IsInsufficientTokens(err):
		if errors.As(err, &ve) && ve.Needed != nil {
			fmt.Println("needs", *ve.Needed, "tokens; free tokens refill at", ve.RefillsAt)
		}
	case errors.As(err, &ve):
		fmt.Println(ve.Status, ve.Code, ve.Message)
	case err != nil:
		log.Fatal(err) // network error or context cancellation
	}
}

func IsNotFound

func IsNotFound(err error) bool

IsNotFound reports whether err is an HTTP 404 from the API.

func IsRateLimited

func IsRateLimited(err error) bool

IsRateLimited reports whether err is an HTTP 429 (too many requests, or a legacy daily limit). Error.RetryAfter says how long to wait when the API knows.

func IsTimeout

func IsTimeout(err error) bool

IsTimeout reports whether a Wait helper gave up while the work was still processing. The work itself continues; you can keep polling.

func NormalizePodcastText

func NormalizePodcastText(s string) string

NormalizePodcastText applies the rule CreatePodcast uses: blank lines are kept as paragraph breaks (one empty line), other runs of spaces and tabs collapse, and the ends are trimmed.

func String

func String(s string) *string

String returns a pointer to s, for optional string fields such as VoiceUpdate.Description.

func TextLength

func TextLength(s string) int

TextLength counts characters the way the API does (UTF-16 code units): letters of every language count one, most emoji count two.

Types

type Account

type Account struct {
	Approved bool `json:"approved"`
	Admin    bool `json:"admin"`
	// Tokens: what generating costs and what you have left.
	Tokens *TokenWallet `json:"tokens,omitempty"`
	// Deprecated: tokens replaced the daily character cap; always nil now.
	// UsedToday still counts characters.
	DailyChars *int64 `json:"dailyChars"`
	UsedToday  int64  `json:"usedToday"`
	// Deprecated: tokens replaced the daily transcription cap; always nil now.
	// AudioUsedToday still counts seconds.
	DailyAudioSec  *int64  `json:"dailyAudioSec"`
	AudioUsedToday float64 `json:"audioUsedToday"`
	// ResetsAt is when the daily counters reset (ISO 8601 UTC).
	ResetsAt string `json:"resetsAt"`
	// MaxClones is the number of voice clones allowed; nil means unlimited.
	MaxClones   *int64 `json:"maxClones"`
	Voices      int64  `json:"voices"`
	Clips       int64  `json:"clips"`
	Chars       int64  `json:"chars"`
	Transcripts int64  `json:"transcripts"`
	// AudioSec is the seconds of audio transcribed, all time.
	AudioSec float64 `json:"audioSec"`
}

Account is the state of your account.

type AudioFormat

type AudioFormat string

AudioFormat is a container format Transcribe accepts. Video files are fine: only the audio track is read.

const (
	AudioWAV  AudioFormat = "wav"
	AudioMP3  AudioFormat = "mp3"
	AudioM4A  AudioFormat = "m4a"
	AudioMP4  AudioFormat = "mp4"
	AudioOGG  AudioFormat = "ogg"
	AudioWebM AudioFormat = "webm"
	AudioFLAC AudioFormat = "flac"
	AudioAAC  AudioFormat = "aac"
)

type Client

type Client struct {
	// contains filtered or unexported fields
}

Client talks to the Voixa API. It is safe for concurrent use.

func NewClient

func NewClient(apiKey string, opts ...Option) *Client

NewClient returns a client authenticated with apiKey. An empty apiKey falls back to the VOIXA_API_KEY environment variable (unless WithIDToken is given).

A client without credentials, or with both an API key and an ID token, is still returned; every call on it fails with a configuration error.

func (*Client) Account

func (c *Client) Account(ctx context.Context) (*Account, error)

Account returns the state of your account: approval, token wallet and counters.

func (*Client) CreatePodcast

func (c *Client) CreatePodcast(ctx context.Context, req PodcastRequest) (*Clip, error)

CreatePodcast turns long text (up to 100 000 characters) into one clip of kind "podcast". It always returns StatusProcessing; poll GetClip or use MakePodcast. Clip.Parts and Clip.PartsDone report progress.

Blank lines are kept as paragraph breaks; other whitespace collapses before the length is checked.

func (*Client) CreateVoice

func (c *Client) CreateVoice(ctx context.Context, req CreateVoiceRequest, wait ...WaitOptions) (*Voice, error)

CreateVoice creates a voice from a reference recording. It uploads the audio, registers the voice, and returns once the voice's sample is ready (or failed: check Voice.Status). Waits with a 2 second interval for up to 3 minutes unless wait says otherwise.

Voice cloning is available for Vietnamese, English, French, German, Italian, Spanish and Portuguese; other languages are rejected with a 400.

Example
package main

import (
	"context"
	"log"
	"os"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	f, err := os.Open("reference.wav") // 3 to 10 seconds of clean speech
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()
	voice, err := vx.CreateVoice(ctx, voixa.CreateVoiceRequest{Name: "Studio narrator", Language: voixa.English, Gender: voixa.GenderFemale, Audio: f})
	if err != nil {
		log.Fatal(err)
	}
	if err := vx.UpdateVoice(ctx, voice.VoiceID, voixa.VoiceUpdate{Shared: voixa.Bool(true)}); err != nil {
		log.Fatal(err)
	}
}

func (*Client) CreateVoiceFromDescription

func (c *Client) CreateVoiceFromDescription(ctx context.Context, req VoiceFromDescriptionRequest, wait ...WaitOptions) (*Voice, error)

CreateVoiceFromDescription goes from a description to a voice in one call: DesignVoice with one option, wait for it, then save it. To pick between several options (the better experience), use DesignVoice with Count 3.

func (*Client) DeleteClips

func (c *Client) DeleteClips(ctx context.Context, clipIDs []string) (int, error)

DeleteClips deletes clips and their audio, and returns how many were deleted. IDs that are not yours are ignored.

func (*Client) DeleteDesign

func (c *Client) DeleteDesign(ctx context.Context, designID string) error

DeleteDesign deletes a design and its draft options. Voices already saved from it are kept.

func (*Client) DeleteTranscripts

func (c *Client) DeleteTranscripts(ctx context.Context, transcriptIDs []string) (int, error)

DeleteTranscripts deletes transcripts with their uploaded audio and result files, and returns how many were deleted. Running jobs are cancelled.

func (*Client) DeleteVoice

func (c *Client) DeleteVoice(ctx context.Context, voiceID string) error

DeleteVoice deletes one of your clones and its reference audio. Clips made with it stay in the library.

func (*Client) DesignVoice

func (c *Client) DesignVoice(ctx context.Context, req DesignVoiceRequest) (*VoiceDesign, error)

DesignVoice generates voice options from a description; no recording needed. It always returns StatusProcessing: each option takes about a minute. Poll GetDesign or use WaitForDesign, then SaveDesign turns the option you like into a voice.

Example
package main

import (
	"context"
	"fmt"
	"log"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	d, err := vx.DesignVoice(ctx, voixa.DesignVoiceRequest{
		Description: "An old man in his seventies, hoarse and deep, speaking slowly",
		Language:    voixa.Vietnamese,
		Count:       3,
	})
	if err != nil {
		log.Fatal(err)
	}
	done, err := vx.WaitForDesign(ctx, d.DesignID) // about a minute per option
	if err != nil {
		log.Fatal(err)
	}
	for _, c := range done.Candidates {
		fmt.Println(c.Index, c.Status, c.URL) // listen, then pick
	}
	voice, err := vx.SaveDesign(ctx, done.DesignID, voixa.SaveDesignRequest{Candidate: 1, Name: "Grandpa Ba", Gender: voixa.GenderMale})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(voice.VoiceID)
}

func (*Client) GetClip

func (c *Client) GetClip(ctx context.Context, clipID string) (*Clip, error)

GetClip returns one clip. While it is ready, Clip.URL is a signed URL valid for one hour; read the clip again for a fresh one.

func (*Client) GetDesign

func (c *Client) GetDesign(ctx context.Context, designID string) (*VoiceDesign, error)

GetDesign returns one design with its options.

func (*Client) GetTranscript

func (c *Client) GetTranscript(ctx context.Context, transcriptID string) (*Transcript, error)

GetTranscript returns one transcript with its text, timed segments and signed download URLs.

func (*Client) GetVoice

func (c *Client) GetVoice(ctx context.Context, voiceID string) (*Voice, error)

GetVoice returns one voice.

func (*Client) Languages

func (c *Client) Languages(ctx context.Context) ([]LanguageInfo, error)

Languages returns the supported languages with labels, the per-call character limit, the sample sentence and the clone-recording guide.

func (*Client) ListClips

func (c *Client) ListClips(ctx context.Context, opts ListClipsOptions) (*ClipPage, error)

ListClips returns your clips, newest first, optionally filtered.

func (*Client) ListDesigns

func (c *Client) ListDesigns(ctx context.Context, opts PageOptions) (*DesignPage, error)

ListDesigns returns your designs, newest first (default 10 per page).

func (*Client) ListTranscripts

func (c *Client) ListTranscripts(ctx context.Context, opts ListTranscriptsOptions) (*TranscriptPage, error)

ListTranscripts returns your transcripts, newest first. List items carry no Text or Segments: read one with GetTranscript for those.

func (*Client) MakePodcast

func (c *Client) MakePodcast(ctx context.Context, req PodcastRequest, wait ...WaitOptions) (*Clip, error)

MakePodcast is CreatePodcast followed by waiting until the episode is no longer processing (default: every 5 seconds for up to 2 hours; a 100 000-character episode takes 20 to 100 minutes depending on the language). Check Clip.Status: a failed podcast is returned, not turned into an error.

Example
package main

import (
	"context"
	"fmt"
	"log"
	"os"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	script, _ := os.ReadFile("episode-12.txt")
	episode, err := vx.MakePodcast(ctx, voixa.PodcastRequest{VoiceID: "vi-truc-ly", Title: "Episode 12", Text: string(script), Project: "my-show"})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(episode.Status, episode.URL, episode.DurationSec)
}

func (*Client) Projects

func (c *Client) Projects(ctx context.Context) ([]Project, error)

Projects returns your content groups with clip and character counts, most recently used first.

func (*Client) STTLanguages

func (c *Client) STTLanguages(ctx context.Context) (*STTInfo, error)

STTLanguages returns the languages you may pin on Transcribe, plus the accepted formats and the limits.

func (*Client) Samples

func (c *Client) Samples(ctx context.Context, language Language) (*Samples, error)

Samples returns the built-in voices with sample audio and the sentence spoken, for a voice picker in your own app. Sample URLs are public files that never change. An empty language returns every language.

func (*Client) SaveDesign

func (c *Client) SaveDesign(ctx context.Context, designID string, req SaveDesignRequest, wait ...WaitOptions) (*Voice, error)

SaveDesign saves one option (Candidate, 0-based) as a voice of your account. From here on it is an ordinary cloned voice: use its VoiceID with Speak, CreatePodcast or Stream. Costs TokenRates.Clone tokens. Waits for the voice's sample like CreateVoice.

func (*Client) Say

func (c *Client) Say(ctx context.Context, req SpeakRequest, wait ...WaitOptions) (*Clip, error)

Say is Speak followed by waiting until the clip is ready (default: every 2 seconds for up to 3 minutes). A failed clip returns an *Error with CodeSynthFailed.

func (*Client) ShareClip

func (c *Client) ShareClip(ctx context.Context, clipID string) (string, error)

ShareClip publishes a ready clip at a public page anyone can open, no account needed, and returns the page's URL.

func (*Client) Speak

func (c *Client) Speak(ctx context.Context, req SpeakRequest) (*Clip, error)

Speak turns text into a clip in one call and returns the clip as the server has it after its wait of about 20 seconds: possibly still StatusProcessing, in which case poll GetClip or use WaitForClip. Say does the waiting for you.

Runs of whitespace and line breaks collapse to one space before the length is checked, exactly as the API counts them.

Example
package main

import (
	"context"
	"fmt"
	"log"
	"time"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	// Get the clip back at once and manage the wait yourself.
	clip, err := vx.Speak(ctx, voixa.SpeakRequest{VoiceID: "vi-truc-ly", Text: "Xin chào!", Wait: voixa.Bool(false)})
	if err != nil {
		log.Fatal(err)
	}
	clip, err = vx.WaitForClip(ctx, clip.ClipID, voixa.WaitOptions{Timeout: 5 * time.Minute})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(clip.Status, clip.URL)
}

func (*Client) Stream

func (c *Client) Stream(ctx context.Context, req StreamRequest) (*SpeechStream, error)

Stream speaks Vietnamese text as it is generated: the first audio usually arrives in about two seconds, and later clauses follow while earlier ones play. Two requests happen under the hood: a one-use ticket from the API (authenticated like every call), then the audio stream itself.

The caller must Close the returned stream. The stream lives as long as ctx: the client's per-call timeout does not cut it off.

Example
package main

import (
	"context"
	"io"
	"log"
	"os"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	s, err := vx.Stream(ctx, voixa.StreamRequest{VoiceID: "vi-truc-ly", Text: "Xin chào, Voixa có thể giúp gì?", SampleRate: 16000})
	if err != nil {
		log.Fatal(err)
	}
	defer s.Close()
	// 16-bit mono PCM at s.SampleRate, clause by clause.
	if _, err := io.Copy(os.Stdout, s); err != nil {
		log.Fatal(err)
	}
}

func (*Client) Transcribe

func (c *Client) Transcribe(ctx context.Context, req TranscribeRequest) (*Transcript, error)

Transcribe uploads audio and starts a transcription: one API call plus a direct upload of the bytes, which starts the job the moment they land.

It returns as soon as the upload finishes, always with StatusProcessing: the work is as long as the audio (about 12 minutes of machine time per hour of speech, spread over several machines) and a cold queue spends two to three minutes starting a machine first. Poll GetTranscript, or call TranscribeFile, which waits for you.

func (*Client) TranscribeFile

func (c *Client) TranscribeFile(ctx context.Context, req TranscribeRequest, wait ...WaitOptions) (*Transcript, error)

TranscribeFile is Transcribe followed by waiting until the transcript is ready (default: every 5 seconds for up to 2 hours). The result carries Text and Segments. A failed transcript returns an *Error with CodeTranscribeFailed.

Example
package main

import (
	"context"
	"fmt"
	"log"
	"os"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	f, err := os.Open("interview.mp3")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()
	t, err := vx.TranscribeFile(ctx, voixa.TranscribeRequest{Audio: f, FileName: "interview.mp3", Title: "Interview"})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(t.Language, t.DurationSec, t.Text)
	fmt.Println(t.SRTURL) // signed for 1 hour
}

func (*Client) UnshareClip

func (c *Client) UnshareClip(ctx context.Context, clipID string) error

UnshareClip withdraws the public page; the link stops working.

func (*Client) UpdateVoice

func (c *Client) UpdateVoice(ctx context.Context, voiceID string, patch VoiceUpdate) error

UpdateVoice renames or describes one of your clones. Shared: true makes it usable by every Voixa account, false withdraws it.

func (*Client) Usage

func (c *Client) Usage(ctx context.Context, opts PageOptions) (*UsagePage, error)

Usage returns the token ledger, newest first: every spend, refund and top-up.

func (*Client) Voices

func (c *Client) Voices(ctx context.Context, language Language) ([]Voice, error)

Voices returns the built-in voices, your own clones and clones other accounts shared. An empty language returns every language.

func (*Client) WaitForClip

func (c *Client) WaitForClip(ctx context.Context, clipID string, wait ...WaitOptions) (*Clip, error)

WaitForClip polls a clip until it is no longer processing (default: every 2 seconds for up to 3 minutes).

func (*Client) WaitForDesign

func (c *Client) WaitForDesign(ctx context.Context, designID string, wait ...WaitOptions) (*VoiceDesign, error)

WaitForDesign polls a design until every option is done (default: every 4 seconds for up to 10 minutes).

func (*Client) WaitForTranscript

func (c *Client) WaitForTranscript(ctx context.Context, transcriptID string, wait ...WaitOptions) (*Transcript, error)

WaitForTranscript polls a transcript until it is no longer processing (default: every 5 seconds for up to 2 hours). DoneSec and DurationSec report progress.

func (*Client) WaitForVoice

func (c *Client) WaitForVoice(ctx context.Context, voiceID string, wait ...WaitOptions) (*Voice, error)

WaitForVoice polls a voice until it is no longer processing (default: every 2 seconds for up to 3 minutes).

type Clip

type Clip struct {
	ClipID    string   `json:"clipId"`
	VoiceID   string   `json:"voiceId"`
	VoiceName string   `json:"voiceName"`
	Language  Language `json:"language"`
	// Text is the text as read, after the API cleaned it.
	Text  string `json:"text"`
	Chars int    `json:"chars"`
	// Project is the content group this clip belongs to, if one was given.
	Project string `json:"project,omitempty"`
	// Source: SourceAPI (API key) or SourceStudio (signed-in web app).
	Source Source `json:"source,omitempty"`
	// Kind is "podcast" for clips made with CreatePodcast; empty for ordinary clips.
	Kind  string `json:"kind,omitempty"`
	Title string `json:"title,omitempty"`
	// Parts of a podcast; PartsDone is reported while the podcast is processing.
	Parts     int `json:"parts,omitempty"`
	PartsDone int `json:"partsDone,omitempty"`
	// Engine is always "voixa-1".
	Engine     string `json:"engine"`
	Format     string `json:"format"`
	SampleRate int    `json:"sampleRate,omitempty"`
	// DurationSec is the length of the audio in seconds.
	DurationSec float64 `json:"durationSec,omitempty"`
	// SynthMs is the synthesis time on the engine, milliseconds.
	SynthMs float64 `json:"synthMs,omitempty"`
	// Tokens charged for this clip. Refunded in full if the clip fails.
	Tokens     int    `json:"tokens,omitempty"`
	Status     Status `json:"status"`
	FailReason string `json:"failReason,omitempty"`
	// URL is a signed WAV URL, valid for 1 hour. Empty while processing.
	URL string `json:"url,omitempty"`
	// ShareURL is the public share page, present after ShareClip. Lives until
	// UnshareClip.
	ShareURL  string `json:"shareUrl,omitempty"`
	CreatedAt string `json:"createdAt"`
	UpdatedAt string `json:"updatedAt"`
}

Clip is one generated audio file: a Speak result or a podcast.

func (*Clip) IsPodcast

func (c *Clip) IsPodcast() bool

IsPodcast reports whether the clip was made with CreatePodcast.

type ClipKind

type ClipKind string

ClipKind filters ListClips: ordinary clips or podcasts.

const (
	ClipKindClip    ClipKind = "clip"
	ClipKindPodcast ClipKind = "podcast"
)

type ClipPage

type ClipPage struct {
	Items  []Clip `json:"items"`
	Cursor string `json:"cursor,omitempty"`
}

ClipPage is one page of ListClips. Pass Cursor back to read the next page; it is empty on the last page.

type CreateVoiceRequest

type CreateVoiceRequest struct {
	Name     string   `json:"name"`
	Language Language `json:"language"`
	// Audio is the reference recording: 3 to 10 seconds of clean speech, WAV or MP3,
	// at most 10 MB.
	Audio io.Reader `json:"-"`
	// Size of Audio in bytes. Optional: taken from the reader when it can tell
	// (*os.File, *bytes.Reader, *bytes.Buffer, *strings.Reader); otherwise the
	// audio is read into memory first.
	Size int64 `json:"-"`
	// Format is "wav" (default) or "mp3".
	Format      AudioFormat `json:"-"`
	Gender      Gender      `json:"gender,omitempty"`
	Description string      `json:"description,omitempty"`
	// RefText is what the speaker says in the recording. Improves Chinese, Japanese
	// and Korean clones a lot; optional elsewhere.
	RefText string `json:"refText,omitempty"`
}

CreateVoiceRequest is the input of CreateVoice.

type DesignCandidate

type DesignCandidate struct {
	Index       int     `json:"index"`
	Status      Status  `json:"status"`
	DurationSec float64 `json:"durationSec,omitempty"`
	// URL is a signed URL of the option (1 hour). Empty while generating, if it
	// failed, or once the design expired.
	URL string `json:"url,omitempty"`
}

DesignCandidate is one generated option of a voice design.

type DesignPage

type DesignPage struct {
	Items  []VoiceDesign `json:"items"`
	Cursor string        `json:"cursor,omitempty"`
}

DesignPage is one page of ListDesigns.

type DesignVoiceRequest

type DesignVoiceRequest struct {
	// Description of the voice in plain words: age, gender, timbre, pace, mood. Any
	// language works; English, Vietnamese and Chinese are understood best. Example:
	// "An old man in his seventies, hoarse and deep, speaking slowly".
	Description string `json:"description"`
	// Language of the voice. Must be one Languages reports as Designable.
	Language Language `json:"language"`
	// Text the options read (max 160 characters, about 8 to 10 seconds; enough to
	// clone from). Defaults to the language's sample sentence.
	Text string `json:"text,omitempty"`
	// Count is how many options to generate, 1 to 3 (default 3). Each costs
	// TokenRates.Design tokens.
	Count int `json:"count,omitempty"`
}

DesignVoiceRequest is the input of DesignVoice.

type Error

type Error struct {
	// Status is the HTTP status (or the status the API would return for a
	// client-side check).
	Status int
	// Message is the human-readable explanation from the API.
	Message string
	// Code is a machine-readable code (see the Code constants), empty when the
	// API did not send one.
	Code string
	// RetryAfter is how long to wait before calling again, from the Retry-After
	// header; zero when absent.
	RetryAfter time.Duration
	// ResetsAt is when a daily counter resets (ISO 8601 UTC), for legacy quota errors.
	ResetsAt string
	// Used and Limit accompany legacy quota errors.
	Used  *int64
	Limit *int64
	// Needed, Available and RefillsAt accompany CodeInsufficientTokens: the tokens
	// the request needs, the tokens you have, and when free tokens refill.
	Needed    *int64
	Available *int64
	RefillsAt string
}

Error is returned for every failed API call, for failed uploads and for the client-side checks that mirror the API's validation. Use errors.As to read it:

var ve *voixa.Error
if errors.As(err, &ve) && ve.Status == 404 { … }

func (*Error) Error

func (e *Error) Error() string

type Gender

type Gender string

Gender of a voice.

const (
	GenderFemale  Gender = "female"
	GenderMale    Gender = "male"
	GenderNeutral Gender = "neutral"
)

type Language

type Language string

Language is a language Voixa speaks. Languages returns labels and clone-recording guides.

const (
	Vietnamese Language = "vi"
	English    Language = "en"
	French     Language = "fr"
	German     Language = "de"
	Italian    Language = "it"
	Spanish    Language = "es"
	Portuguese Language = "pt"
	Chinese    Language = "zh"
	Japanese   Language = "ja"
	Korean     Language = "ko"
)

Languages Voixa speaks.

type LanguageInfo

type LanguageInfo struct {
	Code  Language `json:"code"`
	Label string   `json:"label"`
	// MaxChars per Speak call in this language.
	MaxChars int `json:"maxChars"`
	// Cloneable is false for languages without voice cloning (Chinese, Japanese,
	// Korean).
	Cloneable bool `json:"cloneable"`
	// Designable is true where DesignVoice works.
	Designable bool `json:"designable"`
	// SampleText is the sentence used for voice samples.
	SampleText string `json:"sampleText"`
	// Clone is the guide for recording a clone reference.
	Clone struct {
		Hint   string `json:"hint"`
		Script string `json:"script"`
	} `json:"clone"`
}

LanguageInfo describes one language Voixa speaks.

type ListClipsOptions

type ListClipsOptions struct {
	VoiceID string
	Project string
	Source  Source
	Kind    ClipKind
	Limit   int
	Cursor  string
}

ListClipsOptions filter ListClips. Zero values mean no filter.

type ListTranscriptsOptions

type ListTranscriptsOptions struct {
	Project string
	Source  Source
	Status  Status
	Limit   int
	Cursor  string
}

ListTranscriptsOptions filter ListTranscripts. Zero values mean no filter.

type Option

type Option func(*Client)

Option configures a Client.

func WithBaseURL

func WithBaseURL(u string) Option

WithBaseURL points the client at another deployment of the API.

func WithHTTPClient

func WithHTTPClient(h *http.Client) Option

WithHTTPClient sets the HTTP client used for every request, including uploads and speech streams. Do not give it a short overall Timeout if you use Stream: the timeout would cut long streams off. Use WithTimeout for API calls instead.

func WithIDToken

func WithIDToken(token string) Option

WithIDToken authenticates as a signed-in Studio user with their ID token (not the access token) instead of an API key. Pass an empty API key to NewClient with it.

func WithTimeout

func WithTimeout(d time.Duration) Option

WithTimeout sets the timeout of each API call (default 30 seconds). Uploads and stream bodies are bounded only by the context you pass.

func WithUserAgent

func WithUserAgent(ua string) Option

WithUserAgent appends a product token to the User-Agent header, for example "my-app/1.2".

type PageOptions

type PageOptions struct {
	Limit  int
	Cursor string
}

PageOptions page through a list. Zero values use the API defaults.

type PodcastRequest

type PodcastRequest struct {
	VoiceID string `json:"voiceId"`
	// Title is shown in the library and on the public page; defaults to the first
	// words of the text.
	Title string `json:"title,omitempty"`
	// Text, up to 100 000 characters (about 80 minutes). Blank lines separate
	// paragraphs (a 600 ms pause); other whitespace collapses. Voixa splits the text
	// into sentences, produces the parts in parallel and joins them into one WAV.
	Text    string `json:"text"`
	Project string `json:"project,omitempty"`
}

PodcastRequest is the input of CreatePodcast and MakePodcast.

type Project

type Project struct {
	Project   string `json:"project"`
	Clips     int    `json:"clips"`
	Chars     int    `json:"chars"`
	UpdatedAt string `json:"updatedAt,omitempty"`
}

Project is a content group with its clip and character counts.

type STTInfo

type STTInfo struct {
	Languages []struct {
		Code  string `json:"code"`
		Label string `json:"label"`
	} `json:"languages"`
	Formats    []AudioFormat    `json:"formats"`
	MaxBytes   int64            `json:"maxBytes"`
	MaxSeconds int64            `json:"maxSeconds"`
	Tasks      []TranscribeTask `json:"tasks"`
	Engine     string           `json:"engine"`
}

STTInfo lists the languages you may pin on Transcribe, the accepted formats and the limits.

type SampleVoice

type SampleVoice struct {
	VoiceID     string    `json:"voiceId"`
	Language    Language  `json:"language"`
	Name        string    `json:"name"`
	Gender      Gender    `json:"gender"`
	Styles      []string  `json:"styles"`
	Kind        VoiceKind `json:"kind"`
	Description string    `json:"description,omitempty"`
	// SampleURL is a public URL of a short WAV sample (cached for a year; the file
	// never changes).
	SampleURL string `json:"sampleUrl"`
	// SampleText is the sentence spoken in the sample.
	SampleText string `json:"sampleText"`
	Engine     string `json:"engine"`
}

SampleVoice is a built-in voice with a public sample, for a voice picker in your own app.

type Samples

type Samples struct {
	Languages []struct {
		Code  Language `json:"code"`
		Label string   `json:"label"`
	} `json:"languages"`
	SampleText map[Language]string `json:"sampleText"`
	Voices     []SampleVoice       `json:"voices"`
}

Samples is the answer of Samples.

type SaveDesignRequest

type SaveDesignRequest struct {
	// Candidate is the 0-based index of the option.
	Candidate   int    `json:"candidate"`
	Name        string `json:"name"`
	Gender      Gender `json:"gender,omitempty"`
	Description string `json:"description,omitempty"`
}

SaveDesignRequest picks one option of a design to keep as a voice.

type Source

type Source string

Source tells where something was created: with an API key or in Studio (the signed-in web app).

const (
	SourceStudio Source = "studio"
	SourceAPI    Source = "api"
)

type SpeakRequest

type SpeakRequest struct {
	VoiceID string `json:"voiceId"`
	// Text, up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Line
	// breaks and repeated spaces collapse to one space before counting.
	Text string `json:"text"`
	// Project is an optional content group, e.g. "episode-12" or "promo Q4". Clips
	// with the same project can be listed together and counted (Projects). Letters,
	// digits, spaces, ". _ -"; max 60.
	Project string `json:"project,omitempty"`
	// Wait up to about 20 seconds on the server for the clip. Nil means true; use
	// Bool(false) to get the clip back at once, still processing.
	Wait *bool `json:"wait,omitempty"`
}

SpeakRequest is the input of Speak and Say.

type SpeechStream

type SpeechStream struct {
	io.ReadCloser
	// SampleRate of the audio, in Hz.
	SampleRate int
	// Format of the body: StreamPCM or StreamWAV.
	Format StreamFormat
	// Tokens charged when the stream opened. Refunded if generation fails.
	Tokens int
}

SpeechStream is Vietnamese speech delivered while it is being voiced. It is an io.ReadCloser over the audio: 16-bit little-endian mono PCM (or the same PCM behind a WAV header), clause by clause. Read it until io.EOF and Close it.

type Status

type Status string

Status is the state of a clip, voice, design or transcript.

const (
	StatusProcessing Status = "processing"
	StatusReady      Status = "ready"
	StatusFailed     Status = "failed"
)

type StreamFormat

type StreamFormat string

StreamFormat is the body format of a speech stream.

const (
	// StreamPCM is raw 16-bit little-endian mono PCM (the default).
	StreamPCM StreamFormat = "pcm"
	// StreamWAV is the same PCM behind a streaming WAV header.
	StreamWAV StreamFormat = "wav"
)

type StreamRequest

type StreamRequest struct {
	// VoiceID of a Vietnamese voice (preset or your clone). For other languages use
	// Speak.
	VoiceID string `json:"voiceId"`
	// Text, up to 5 000 characters. Cleaned and counted like Speak.
	Text string `json:"text"`
	// SampleRate of the output: 48000, 24000 (default), 16000 or 8000. 8000 and
	// 16000 suit telephony.
	SampleRate int `json:"sampleRate,omitempty"`
	// Format: StreamPCM (default) or StreamWAV.
	Format StreamFormat `json:"format,omitempty"`
}

StreamRequest is the input of Stream.

type TokenRates

type TokenRates struct {
	// ViFast per character, Vietnamese.
	ViFast float64 `json:"viFast"`
	// Standard per character: English, French, German, Italian, Spanish, Portuguese.
	Standard float64 `json:"standard"`
	// CJK per character: Chinese, Japanese, Korean.
	CJK float64 `json:"cjk"`
	// STTPerSec per second of transcribed audio.
	STTPerSec float64 `json:"sttPerSec"`
	// Clone per voice clone created (also charged when you save a designed voice).
	Clone float64 `json:"clone"`
	// Design per option generated by DesignVoice (3 options cost 3 times this).
	Design float64 `json:"design"`
	// Stream per character, Vietnamese streaming.
	Stream float64 `json:"stream"`
}

TokenRates say how many tokens one unit of work costs. One token is one Vietnamese character read on the fast path.

func (TokenRates) EstimateClone

func (r TokenRates) EstimateClone() int

EstimateClone is the cost of creating a voice clone or saving a designed voice.

func (TokenRates) EstimateDesign

func (r TokenRates) EstimateDesign(options int) int

EstimateDesign is the cost of a voice design with the given number of options.

func (TokenRates) EstimateSpeak

func (r TokenRates) EstimateSpeak(language Language, chars int) int

EstimateSpeak is the cost of reading chars characters in language, for Speak or CreatePodcast.

Example
package main

import (
	"context"
	"fmt"
	"log"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("")

	acc, err := vx.Account(ctx)
	if err != nil {
		log.Fatal(err)
	}
	if acc.Tokens != nil {
		fmt.Println(acc.Tokens.Rates.EstimateSpeak(voixa.Vietnamese, 1200))
	}
}

func (TokenRates) EstimateStream

func (r TokenRates) EstimateStream(chars int) int

EstimateStream is the cost of streaming chars characters.

func (TokenRates) EstimateTranscription

func (r TokenRates) EstimateTranscription(seconds float64) int

EstimateTranscription is the cost of transcribing seconds of audio. Transcription is charged per second of audio after it finishes.

type TokenWallet

type TokenWallet struct {
	// Available is spendable right now (free today + balance). Nil for accounts
	// that are not limited.
	Available *int64 `json:"available"`
	// FreeToday is the free tokens left today. They refill to DailyFree at
	// RefillsAt; unused free tokens do not carry over.
	FreeToday int64 `json:"freeToday"`
	DailyFree int64 `json:"dailyFree"`
	// Balance is tokens added to the account, spent after today's free tokens. Can be
	// slightly negative after a long transcription.
	Balance int64 `json:"balance"`
	// UsedToday is tokens spent today, net of refunds.
	UsedToday int64 `json:"usedToday"`
	// RefillsAt is when free tokens refill (ISO 8601 UTC).
	RefillsAt string     `json:"refillsAt"`
	Rates     TokenRates `json:"rates"`
	Unlimited bool       `json:"unlimited"`
}

TokenWallet is a free allowance that refills daily plus a balance that carries over.

type TranscribeRequest

type TranscribeRequest struct {
	// Audio is the audio (or video) file.
	Audio io.Reader
	// Size of Audio in bytes. Optional: taken from the reader when it can tell
	// (*os.File, *bytes.Reader, *bytes.Buffer, *strings.Reader); otherwise the
	// audio is read into memory first.
	Size int64
	// Format of the container. Inferred from FileName (or the name of an *os.File)
	// when empty; MP3 otherwise.
	Format AudioFormat
	// FileName is the original file name: shown in the library, not used as a key.
	FileName string
	Title    string
	// Language pins the spoken language ("vi", "en", "ja", …). Leave it empty to let
	// Voixa detect it, which it does well; see STTLanguages for the list.
	Language string
	// Task: TaskTranslate recognises the speech and returns English in one pass.
	Task TranscribeTask
	// Words asks for per-word timestamps (about 10% slower).
	Words bool
	// Prompt gives spelling hints for names and jargon, read as if it preceded the audio.
	Prompt  string
	Project string
}

TranscribeRequest is the input of Transcribe and TranscribeFile.

type TranscribeTask

type TranscribeTask string

TranscribeTask: TaskTranscribe keeps the spoken language; TaskTranslate recognises the speech and returns English in one pass.

const (
	TaskTranscribe TranscribeTask = "transcribe"
	TaskTranslate  TranscribeTask = "translate"
)

type Transcript

type Transcript struct {
	TranscriptID string `json:"transcriptId"`
	Title        string `json:"title,omitempty"`
	// FileName is the file name you passed, for your own bookkeeping.
	FileName string      `json:"fileName,omitempty"`
	Format   AudioFormat `json:"format"`
	// RequestedLanguage is the language you asked for, if you pinned one.
	RequestedLanguage string `json:"requestedLanguage,omitempty"`
	// Language of the speech: detected when you did not pin one.
	Language string `json:"language,omitempty"`
	// LanguageProb is the confidence of the language detection, 0-1.
	LanguageProb float64        `json:"languageProb,omitempty"`
	Task         TranscribeTask `json:"task"`
	Words        bool           `json:"words"`
	Source       Source         `json:"source,omitempty"`
	Project      string         `json:"project,omitempty"`
	// Engine is always "voixa-stt-1".
	Engine     string `json:"engine"`
	AudioBytes int64  `json:"audioBytes"`
	// DurationSec is the length of the audio (known once the file is decoded).
	DurationSec float64 `json:"durationSec,omitempty"`
	// Chars of transcript text.
	Chars        int `json:"chars,omitempty"`
	SegmentCount int `json:"segmentCount,omitempty"`
	// TranscribeMs is machine time spent recognising, summed over all parts, so it
	// exceeds the time you waited.
	TranscribeMs float64 `json:"transcribeMs,omitempty"`
	// RTF is machine time divided by audio length. Below 1 is faster than real time
	// on one machine.
	RTF float64 `json:"rtf,omitempty"`
	// Parts: long files are recognised in parallel on several machines and joined
	// into one transcript at the end.
	Parts int `json:"parts,omitempty"`
	// PartsDone so far, while processing.
	PartsDone int `json:"partsDone,omitempty"`
	// DoneSec is the seconds of audio already recognised, while processing.
	DoneSec float64 `json:"doneSec,omitempty"`
	// JobStatus is what the job is doing while processing: RUNNABLE (waiting for a
	// machine), STARTING, RUNNING. A cold queue takes two to three minutes to reach
	// RUNNING.
	JobStatus  string `json:"jobStatus,omitempty"`
	Status     Status `json:"status"`
	FailReason string `json:"failReason,omitempty"`
	// AudioURL is a signed URL of the audio you uploaded, valid for 1 hour.
	AudioURL string `json:"audioUrl,omitempty"`
	// Signed URLs of the results, valid for 1 hour. Only when you read one
	// transcript (GetTranscript), not in lists.
	JSONURL string `json:"jsonUrl,omitempty"`
	SRTURL  string `json:"srtUrl,omitempty"`
	VTTURL  string `json:"vttUrl,omitempty"`
	TXTURL  string `json:"txtUrl,omitempty"`
	// Text is the full text. Only when you read one transcript, not in lists.
	Text string `json:"text,omitempty"`
	// Segments are the timed segments. Only when you read one transcript.
	Segments  []TranscriptSegment `json:"segments,omitempty"`
	CreatedAt string              `json:"createdAt"`
	UpdatedAt string              `json:"updatedAt"`
}

Transcript is the result of a transcription.

type TranscriptPage

type TranscriptPage struct {
	Items  []Transcript `json:"items"`
	Cursor string       `json:"cursor,omitempty"`
}

TranscriptPage is one page of ListTranscripts.

type TranscriptSegment

type TranscriptSegment struct {
	ID int `json:"id"`
	// Start and End are seconds from the start of the audio.
	Start float64 `json:"start"`
	End   float64 `json:"end"`
	Text  string  `json:"text"`
	// Words is set only when the request asked for Words.
	Words []TranscriptWord `json:"words,omitempty"`
}

TranscriptSegment is a timed piece of the transcript.

type TranscriptWord

type TranscriptWord struct {
	Start float64 `json:"start"`
	End   float64 `json:"end"`
	Word  string  `json:"word"`
	// Prob is the confidence, 0-1.
	Prob float64 `json:"prob"`
}

TranscriptWord is one word with its timing, when the request asked for Words.

type UsageEntry

type UsageEntry struct {
	ID string `json:"id"`
	// Kind: speak, podcast, clone, design, transcribe, stream, refund or grant.
	Kind   string `json:"kind"`
	Tokens int64  `json:"tokens"`
	// Free and Paid split Tokens between today's free tokens and the balance.
	Free int64 `json:"free"`
	Paid int64 `json:"paid"`
	// Ref is the clipId, voiceId or transcriptId concerned.
	Ref  string `json:"ref,omitempty"`
	Note string `json:"note,omitempty"`
	At   string `json:"at"`
}

UsageEntry is one line of the token ledger: negative for spends, positive for refunds and top-ups.

type UsagePage

type UsagePage struct {
	Items  []UsageEntry `json:"items"`
	Cursor string       `json:"cursor,omitempty"`
}

UsagePage is one page of Usage.

type Voice

type Voice struct {
	VoiceID string `json:"voiceId"`
	// Kind is VoiceKindPreset for a built-in voice, VoiceKindClone for one created
	// from a reference recording (yours, or a built-in one).
	Kind VoiceKind `json:"kind"`
	// Mine is true for voices your account created: those can be renamed, shared
	// and deleted.
	Mine bool `json:"mine"`
	// Shared: a clone its owner shared with every account (or your own clone, if
	// you shared it).
	Shared      bool     `json:"shared"`
	Language    Language `json:"language"`
	Name        string   `json:"name"`
	Gender      Gender   `json:"gender"`
	Styles      []string `json:"styles"`
	Description string   `json:"description,omitempty"`
	// SampleURL is a short sample of the voice. Signed URL for clones (1 hour);
	// empty while a clone is still processing.
	SampleURL  string `json:"sampleUrl,omitempty"`
	Status     Status `json:"status"`
	FailReason string `json:"failReason,omitempty"`
	// Engine is always "voixa-1".
	Engine string `json:"engine"`
	// DesignID is set when the voice was saved from a voice design (SaveDesign).
	DesignID  string `json:"designId,omitempty"`
	CreatedAt string `json:"createdAt"`
	UpdatedAt string `json:"updatedAt"`
}

Voice is a built-in voice or a clone.

type VoiceDesign

type VoiceDesign struct {
	DesignID    string   `json:"designId"`
	Language    Language `json:"language"`
	Description string   `json:"description"`
	// Instruct is what the voice engine was given: your description, translated to
	// English when needed.
	Instruct   string            `json:"instruct,omitempty"`
	Text       string            `json:"text"`
	Count      int               `json:"count"`
	Status     Status            `json:"status"`
	FailReason string            `json:"failReason,omitempty"`
	Candidates []DesignCandidate `json:"candidates"`
	// SavedVoiceIDs are the voices saved from this design.
	SavedVoiceIDs []string `json:"savedVoiceIds"`
	Source        Source   `json:"source,omitempty"`
	Tokens        int      `json:"tokens,omitempty"`
	ExpiresAt     string   `json:"expiresAt"`
	Expired       bool     `json:"expired"`
	CreatedAt     string   `json:"createdAt"`
	UpdatedAt     string   `json:"updatedAt"`
}

VoiceDesign holds options generated from a description. Options are drafts kept for 23 hours (ExpiresAt); save the one you like with SaveDesign to get a permanent voice.

type VoiceFromDescriptionRequest

type VoiceFromDescriptionRequest struct {
	Description string
	Language    Language
	Text        string
	Name        string
	Gender      Gender
}

VoiceFromDescriptionRequest is the input of CreateVoiceFromDescription.

type VoiceKind

type VoiceKind string

VoiceKind: a built-in preset, or a clone made from a reference recording (yours, or a built-in one).

const (
	VoiceKindPreset VoiceKind = "preset"
	VoiceKindClone  VoiceKind = "clone"
)

type VoiceUpdate

type VoiceUpdate struct {
	Name string
	// Description: a pointer to "" clears it.
	Description *string
	Gender      Gender
	// Styles: a non-nil empty slice clears them.
	Styles []string
	// Shared: true makes the clone usable by every Voixa account, false withdraws it.
	Shared *bool
}

VoiceUpdate changes a clone. Nil or empty fields are left as they are.

func (VoiceUpdate) MarshalJSON

func (u VoiceUpdate) MarshalJSON() ([]byte, error)

MarshalJSON sends only the fields that change.

type WaitOptions

type WaitOptions struct {
	// PollInterval between two status reads.
	PollInterval time.Duration
	// Timeout: give up after this long with an *Error of code CodeTimeout. Cold
	// engines can take about a minute.
	Timeout time.Duration
}

WaitOptions tune the Wait helpers and the methods that wait for you. Zero fields keep the method's default.

Directories

Path Synopsis
cmd
e2e command
Command e2e checks every public method of the Go client against a live Voixa API.
Command e2e checks every public method of the Go client against a live Voixa API.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL