Documentation
¶
Overview ¶
Package voixa is the official Go client for the Voixa AI speech API.
Text to speech in ten languages (built-in voices, voice cloning from a short sample, voice design from a written description, podcasts up to 100 000 characters, streaming Vietnamese speech) and speech to text: audio or video in, transcript with timestamps, SRT and VTT out.
The package has no dependencies outside the standard library.
Authentication ¶
Machine callers use an API key (create one in Studio, API keys). The key is sent in the x-api-key header. NewClient falls back to the VOIXA_API_KEY environment variable when the key argument is empty. A signed-in web user can be represented with WithIDToken instead.
Two kinds of work ¶
Text to speech (Client.Speak, Client.CreatePodcast) answers quickly on a warm engine, so the API waits a little for you. Speech to text (Client.Transcribe) is as long as the audio, so it never waits: you get StatusProcessing and poll, or let Client.TranscribeFile do it.
How Speak behaves ¶
The API synthesises in the background and waits up to about 20 seconds for the clip. On a warm engine that is a few seconds and the clip comes back ready with a URL. On a cold engine it may come back processing; Client.Say handles that by polling, Client.Speak leaves it to you.
Example ¶
package main
import (
"context"
"fmt"
"log"
"os"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient(os.Getenv("VOIXA_API_KEY")) // create a key in Studio, API keys
voices, err := vx.Voices(ctx, voixa.English)
if err != nil {
log.Fatal(err)
}
clip, err := vx.Say(ctx, voixa.SpeakRequest{
VoiceID: voices[0].VoiceID,
Text: "The lighthouse keeper switched the lamp on at dusk.",
Project: "episode-12", // optional: group clips by content
})
if err != nil {
log.Fatal(err)
}
fmt.Println(clip.URL, clip.DurationSec) // signed WAV URL, valid for 1 hour
}
Output:
Index ¶
- Constants
- func Bool(b bool) *bool
- func CollapseSpace(s string) string
- func ErrorCode(err error) string
- func ErrorStatus(err error) int
- func IsInsufficientTokens(err error) bool
- func IsNotFound(err error) bool
- func IsRateLimited(err error) bool
- func IsTimeout(err error) bool
- func NormalizePodcastText(s string) string
- func String(s string) *string
- func TextLength(s string) int
- type Account
- type AudioFormat
- type Client
- func (c *Client) Account(ctx context.Context) (*Account, error)
- func (c *Client) CreatePodcast(ctx context.Context, req PodcastRequest) (*Clip, error)
- func (c *Client) CreateVoice(ctx context.Context, req CreateVoiceRequest, wait ...WaitOptions) (*Voice, error)
- func (c *Client) CreateVoiceFromDescription(ctx context.Context, req VoiceFromDescriptionRequest, wait ...WaitOptions) (*Voice, error)
- func (c *Client) DeleteClips(ctx context.Context, clipIDs []string) (int, error)
- func (c *Client) DeleteDesign(ctx context.Context, designID string) error
- func (c *Client) DeleteTranscripts(ctx context.Context, transcriptIDs []string) (int, error)
- func (c *Client) DeleteVoice(ctx context.Context, voiceID string) error
- func (c *Client) DesignVoice(ctx context.Context, req DesignVoiceRequest) (*VoiceDesign, error)
- func (c *Client) GetClip(ctx context.Context, clipID string) (*Clip, error)
- func (c *Client) GetDesign(ctx context.Context, designID string) (*VoiceDesign, error)
- func (c *Client) GetTranscript(ctx context.Context, transcriptID string) (*Transcript, error)
- func (c *Client) GetVoice(ctx context.Context, voiceID string) (*Voice, error)
- func (c *Client) Languages(ctx context.Context) ([]LanguageInfo, error)
- func (c *Client) ListClips(ctx context.Context, opts ListClipsOptions) (*ClipPage, error)
- func (c *Client) ListDesigns(ctx context.Context, opts PageOptions) (*DesignPage, error)
- func (c *Client) ListTranscripts(ctx context.Context, opts ListTranscriptsOptions) (*TranscriptPage, error)
- func (c *Client) MakePodcast(ctx context.Context, req PodcastRequest, wait ...WaitOptions) (*Clip, error)
- func (c *Client) Projects(ctx context.Context) ([]Project, error)
- func (c *Client) STTLanguages(ctx context.Context) (*STTInfo, error)
- func (c *Client) Samples(ctx context.Context, language Language) (*Samples, error)
- func (c *Client) SaveDesign(ctx context.Context, designID string, req SaveDesignRequest, ...) (*Voice, error)
- func (c *Client) Say(ctx context.Context, req SpeakRequest, wait ...WaitOptions) (*Clip, error)
- func (c *Client) ShareClip(ctx context.Context, clipID string) (string, error)
- func (c *Client) Speak(ctx context.Context, req SpeakRequest) (*Clip, error)
- func (c *Client) Stream(ctx context.Context, req StreamRequest) (*SpeechStream, error)
- func (c *Client) Transcribe(ctx context.Context, req TranscribeRequest) (*Transcript, error)
- func (c *Client) TranscribeFile(ctx context.Context, req TranscribeRequest, wait ...WaitOptions) (*Transcript, error)
- func (c *Client) UnshareClip(ctx context.Context, clipID string) error
- func (c *Client) UpdateVoice(ctx context.Context, voiceID string, patch VoiceUpdate) error
- func (c *Client) Usage(ctx context.Context, opts PageOptions) (*UsagePage, error)
- func (c *Client) Voices(ctx context.Context, language Language) ([]Voice, error)
- func (c *Client) WaitForClip(ctx context.Context, clipID string, wait ...WaitOptions) (*Clip, error)
- func (c *Client) WaitForDesign(ctx context.Context, designID string, wait ...WaitOptions) (*VoiceDesign, error)
- func (c *Client) WaitForTranscript(ctx context.Context, transcriptID string, wait ...WaitOptions) (*Transcript, error)
- func (c *Client) WaitForVoice(ctx context.Context, voiceID string, wait ...WaitOptions) (*Voice, error)
- type Clip
- type ClipKind
- type ClipPage
- type CreateVoiceRequest
- type DesignCandidate
- type DesignPage
- type DesignVoiceRequest
- type Error
- type Gender
- type Language
- type LanguageInfo
- type ListClipsOptions
- type ListTranscriptsOptions
- type Option
- type PageOptions
- type PodcastRequest
- type Project
- type STTInfo
- type SampleVoice
- type Samples
- type SaveDesignRequest
- type Source
- type SpeakRequest
- type SpeechStream
- type Status
- type StreamFormat
- type StreamRequest
- type TokenRates
- type TokenWallet
- type TranscribeRequest
- type TranscribeTask
- type Transcript
- type TranscriptPage
- type TranscriptSegment
- type TranscriptWord
- type UsageEntry
- type UsagePage
- type Voice
- type VoiceDesign
- type VoiceFromDescriptionRequest
- type VoiceKind
- type VoiceUpdate
- type WaitOptions
Examples ¶
Constants ¶
const ( // MaxSpeakChars is the hard limit for one Speak call. Chinese, Japanese and // Korean have a lower limit per call (see Languages). MaxSpeakChars = 5_000 // MaxPodcastChars is the character limit of one podcast. MaxPodcastChars = 100_000 // MaxAudioBytes is the largest file Transcribe accepts: the most a single // upload can carry. MaxAudioBytes int64 = 5 * 1024 * 1024 * 1024 // MaxAudioMinutes is the longest audio one transcription accepts. MaxAudioMinutes = 480 )
Limits enforced by the API, checked locally before a request is sent.
const ( // CodeNotApproved: the account has not been approved yet (403). CodeNotApproved = "NOT_APPROVED" // CodeInsufficientTokens: not enough tokens for this request (402). See // Error.Needed, Error.Available and Error.RefillsAt. CodeInsufficientTokens = "INSUFFICIENT_TOKENS" // CodeQuota: legacy daily character limit reached (429). CodeQuota = "QUOTA" // CodeCloneLimit: voice clone limit reached (403). CodeCloneLimit = "CLONE_LIMIT" // CodeVoiceNotReady: the voice is still processing or failed (409). CodeVoiceNotReady = "VOICE_NOT_READY" // CodeInvalid: request validation failed (400). CodeInvalid = "INVALID" // CodeAudioQuota: legacy daily transcription limit reached (429). CodeAudioQuota = "AUDIO_QUOTA" // CodeAudioTooLarge: the audio file is larger than one upload can carry (400). CodeAudioTooLarge = "AUDIO_TOO_LARGE" // CodeNoAudio: nothing was uploaded for this transcript (400). CodeNoAudio = "NO_AUDIO" // CodeTimeout: a Wait helper gave up while the work was still processing (504). CodeTimeout = "TIMEOUT" // CodeSynthFailed: Say finished with a failed clip (502). CodeSynthFailed = "SYNTH_FAILED" // CodeTranscribeFailed: TranscribeFile finished with a failed transcript (502). CodeTranscribeFailed = "TRANSCRIBE_FAILED" // CodeDesignFailed: CreateVoiceFromDescription finished with a failed design (502). CodeDesignFailed = "DESIGN_FAILED" )
Machine-readable error codes that accompany the message in Error.Code.
const DefaultBaseURL = "https://api.voixa.vovix.io/v1"
DefaultBaseURL is the public Voixa API.
const DefaultTimeout = 30 * time.Second
DefaultTimeout is the per-request timeout for API calls. It does not apply to audio uploads or to the body of a speech stream, which last as long as they need.
const Version = "0.2.0"
Version is the version of this client, sent in the User-Agent header.
Variables ¶
This section is empty.
Functions ¶
func CollapseSpace ¶
CollapseSpace applies the rule Speak and Stream use before counting: every run of whitespace, line breaks included, becomes one space, and the ends are trimmed.
func ErrorStatus ¶
ErrorStatus returns the HTTP status of a *Error anywhere in err's chain, or 0.
func IsInsufficientTokens ¶
IsInsufficientTokens reports whether err means the account does not have enough tokens for the request (HTTP 402).
Example ¶
package main
import (
"context"
"errors"
"fmt"
"log"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
_, err := vx.Say(ctx, voixa.SpeakRequest{VoiceID: "vi-truc-ly", Text: "Xin chào!"})
var ve *voixa.Error
switch {
case voixa.IsInsufficientTokens(err):
if errors.As(err, &ve) && ve.Needed != nil {
fmt.Println("needs", *ve.Needed, "tokens; free tokens refill at", ve.RefillsAt)
}
case errors.As(err, &ve):
fmt.Println(ve.Status, ve.Code, ve.Message)
case err != nil:
log.Fatal(err) // network error or context cancellation
}
}
Output:
func IsNotFound ¶
IsNotFound reports whether err is an HTTP 404 from the API.
func IsRateLimited ¶
IsRateLimited reports whether err is an HTTP 429 (too many requests, or a legacy daily limit). Error.RetryAfter says how long to wait when the API knows.
func IsTimeout ¶
IsTimeout reports whether a Wait helper gave up while the work was still processing. The work itself continues; you can keep polling.
func NormalizePodcastText ¶
NormalizePodcastText applies the rule CreatePodcast uses: blank lines are kept as paragraph breaks (one empty line), other runs of spaces and tabs collapse, and the ends are trimmed.
func String ¶
String returns a pointer to s, for optional string fields such as VoiceUpdate.Description.
func TextLength ¶
TextLength counts characters the way the API does (UTF-16 code units): letters of every language count one, most emoji count two.
Types ¶
type Account ¶
type Account struct {
Approved bool `json:"approved"`
Admin bool `json:"admin"`
// Tokens: what generating costs and what you have left.
Tokens *TokenWallet `json:"tokens,omitempty"`
// Deprecated: tokens replaced the daily character cap; always nil now.
// UsedToday still counts characters.
DailyChars *int64 `json:"dailyChars"`
UsedToday int64 `json:"usedToday"`
// Deprecated: tokens replaced the daily transcription cap; always nil now.
// AudioUsedToday still counts seconds.
DailyAudioSec *int64 `json:"dailyAudioSec"`
AudioUsedToday float64 `json:"audioUsedToday"`
// ResetsAt is when the daily counters reset (ISO 8601 UTC).
ResetsAt string `json:"resetsAt"`
// MaxClones is the number of voice clones allowed; nil means unlimited.
MaxClones *int64 `json:"maxClones"`
Voices int64 `json:"voices"`
Clips int64 `json:"clips"`
Chars int64 `json:"chars"`
Transcripts int64 `json:"transcripts"`
// AudioSec is the seconds of audio transcribed, all time.
AudioSec float64 `json:"audioSec"`
}
Account is the state of your account.
type AudioFormat ¶
type AudioFormat string
AudioFormat is a container format Transcribe accepts. Video files are fine: only the audio track is read.
const ( AudioWAV AudioFormat = "wav" AudioMP3 AudioFormat = "mp3" AudioM4A AudioFormat = "m4a" AudioMP4 AudioFormat = "mp4" AudioOGG AudioFormat = "ogg" AudioWebM AudioFormat = "webm" AudioFLAC AudioFormat = "flac" AudioAAC AudioFormat = "aac" )
type Client ¶
type Client struct {
// contains filtered or unexported fields
}
Client talks to the Voixa API. It is safe for concurrent use.
func NewClient ¶
NewClient returns a client authenticated with apiKey. An empty apiKey falls back to the VOIXA_API_KEY environment variable (unless WithIDToken is given).
A client without credentials, or with both an API key and an ID token, is still returned; every call on it fails with a configuration error.
func (*Client) Account ¶
Account returns the state of your account: approval, token wallet and counters.
func (*Client) CreatePodcast ¶
CreatePodcast turns long text (up to 100 000 characters) into one clip of kind "podcast". It always returns StatusProcessing; poll GetClip or use MakePodcast. Clip.Parts and Clip.PartsDone report progress.
Blank lines are kept as paragraph breaks; other whitespace collapses before the length is checked.
func (*Client) CreateVoice ¶
func (c *Client) CreateVoice(ctx context.Context, req CreateVoiceRequest, wait ...WaitOptions) (*Voice, error)
CreateVoice creates a voice from a reference recording. It uploads the audio, registers the voice, and returns once the voice's sample is ready (or failed: check Voice.Status). Waits with a 2 second interval for up to 3 minutes unless wait says otherwise.
Voice cloning is available for Vietnamese, English, French, German, Italian, Spanish and Portuguese; other languages are rejected with a 400.
Example ¶
package main
import (
"context"
"log"
"os"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
f, err := os.Open("reference.wav") // 3 to 10 seconds of clean speech
if err != nil {
log.Fatal(err)
}
defer f.Close()
voice, err := vx.CreateVoice(ctx, voixa.CreateVoiceRequest{Name: "Studio narrator", Language: voixa.English, Gender: voixa.GenderFemale, Audio: f})
if err != nil {
log.Fatal(err)
}
if err := vx.UpdateVoice(ctx, voice.VoiceID, voixa.VoiceUpdate{Shared: voixa.Bool(true)}); err != nil {
log.Fatal(err)
}
}
Output:
func (*Client) CreateVoiceFromDescription ¶
func (c *Client) CreateVoiceFromDescription(ctx context.Context, req VoiceFromDescriptionRequest, wait ...WaitOptions) (*Voice, error)
CreateVoiceFromDescription goes from a description to a voice in one call: DesignVoice with one option, wait for it, then save it. To pick between several options (the better experience), use DesignVoice with Count 3.
func (*Client) DeleteClips ¶
DeleteClips deletes clips and their audio, and returns how many were deleted. IDs that are not yours are ignored.
func (*Client) DeleteDesign ¶
DeleteDesign deletes a design and its draft options. Voices already saved from it are kept.
func (*Client) DeleteTranscripts ¶
DeleteTranscripts deletes transcripts with their uploaded audio and result files, and returns how many were deleted. Running jobs are cancelled.
func (*Client) DeleteVoice ¶
DeleteVoice deletes one of your clones and its reference audio. Clips made with it stay in the library.
func (*Client) DesignVoice ¶
func (c *Client) DesignVoice(ctx context.Context, req DesignVoiceRequest) (*VoiceDesign, error)
DesignVoice generates voice options from a description; no recording needed. It always returns StatusProcessing: each option takes about a minute. Poll GetDesign or use WaitForDesign, then SaveDesign turns the option you like into a voice.
Example ¶
package main
import (
"context"
"fmt"
"log"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
d, err := vx.DesignVoice(ctx, voixa.DesignVoiceRequest{
Description: "An old man in his seventies, hoarse and deep, speaking slowly",
Language: voixa.Vietnamese,
Count: 3,
})
if err != nil {
log.Fatal(err)
}
done, err := vx.WaitForDesign(ctx, d.DesignID) // about a minute per option
if err != nil {
log.Fatal(err)
}
for _, c := range done.Candidates {
fmt.Println(c.Index, c.Status, c.URL) // listen, then pick
}
voice, err := vx.SaveDesign(ctx, done.DesignID, voixa.SaveDesignRequest{Candidate: 1, Name: "Grandpa Ba", Gender: voixa.GenderMale})
if err != nil {
log.Fatal(err)
}
fmt.Println(voice.VoiceID)
}
Output:
func (*Client) GetClip ¶
GetClip returns one clip. While it is ready, Clip.URL is a signed URL valid for one hour; read the clip again for a fresh one.
func (*Client) GetTranscript ¶
GetTranscript returns one transcript with its text, timed segments and signed download URLs.
func (*Client) Languages ¶
func (c *Client) Languages(ctx context.Context) ([]LanguageInfo, error)
Languages returns the supported languages with labels, the per-call character limit, the sample sentence and the clone-recording guide.
func (*Client) ListDesigns ¶
func (c *Client) ListDesigns(ctx context.Context, opts PageOptions) (*DesignPage, error)
ListDesigns returns your designs, newest first (default 10 per page).
func (*Client) ListTranscripts ¶
func (c *Client) ListTranscripts(ctx context.Context, opts ListTranscriptsOptions) (*TranscriptPage, error)
ListTranscripts returns your transcripts, newest first. List items carry no Text or Segments: read one with GetTranscript for those.
func (*Client) MakePodcast ¶
func (c *Client) MakePodcast(ctx context.Context, req PodcastRequest, wait ...WaitOptions) (*Clip, error)
MakePodcast is CreatePodcast followed by waiting until the episode is no longer processing (default: every 5 seconds for up to 2 hours; a 100 000-character episode takes 20 to 100 minutes depending on the language). Check Clip.Status: a failed podcast is returned, not turned into an error.
Example ¶
package main
import (
"context"
"fmt"
"log"
"os"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
script, _ := os.ReadFile("episode-12.txt")
episode, err := vx.MakePodcast(ctx, voixa.PodcastRequest{VoiceID: "vi-truc-ly", Title: "Episode 12", Text: string(script), Project: "my-show"})
if err != nil {
log.Fatal(err)
}
fmt.Println(episode.Status, episode.URL, episode.DurationSec)
}
Output:
func (*Client) Projects ¶
Projects returns your content groups with clip and character counts, most recently used first.
func (*Client) STTLanguages ¶
STTLanguages returns the languages you may pin on Transcribe, plus the accepted formats and the limits.
func (*Client) Samples ¶
Samples returns the built-in voices with sample audio and the sentence spoken, for a voice picker in your own app. Sample URLs are public files that never change. An empty language returns every language.
func (*Client) SaveDesign ¶
func (c *Client) SaveDesign(ctx context.Context, designID string, req SaveDesignRequest, wait ...WaitOptions) (*Voice, error)
SaveDesign saves one option (Candidate, 0-based) as a voice of your account. From here on it is an ordinary cloned voice: use its VoiceID with Speak, CreatePodcast or Stream. Costs TokenRates.Clone tokens. Waits for the voice's sample like CreateVoice.
func (*Client) Say ¶
func (c *Client) Say(ctx context.Context, req SpeakRequest, wait ...WaitOptions) (*Clip, error)
Say is Speak followed by waiting until the clip is ready (default: every 2 seconds for up to 3 minutes). A failed clip returns an *Error with CodeSynthFailed.
func (*Client) ShareClip ¶
ShareClip publishes a ready clip at a public page anyone can open, no account needed, and returns the page's URL.
func (*Client) Speak ¶
Speak turns text into a clip in one call and returns the clip as the server has it after its wait of about 20 seconds: possibly still StatusProcessing, in which case poll GetClip or use WaitForClip. Say does the waiting for you.
Runs of whitespace and line breaks collapse to one space before the length is checked, exactly as the API counts them.
Example ¶
package main
import (
"context"
"fmt"
"log"
"time"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
// Get the clip back at once and manage the wait yourself.
clip, err := vx.Speak(ctx, voixa.SpeakRequest{VoiceID: "vi-truc-ly", Text: "Xin chào!", Wait: voixa.Bool(false)})
if err != nil {
log.Fatal(err)
}
clip, err = vx.WaitForClip(ctx, clip.ClipID, voixa.WaitOptions{Timeout: 5 * time.Minute})
if err != nil {
log.Fatal(err)
}
fmt.Println(clip.Status, clip.URL)
}
Output:
func (*Client) Stream ¶
func (c *Client) Stream(ctx context.Context, req StreamRequest) (*SpeechStream, error)
Stream speaks Vietnamese text as it is generated: the first audio usually arrives in about two seconds, and later clauses follow while earlier ones play. Two requests happen under the hood: a one-use ticket from the API (authenticated like every call), then the audio stream itself.
The caller must Close the returned stream. The stream lives as long as ctx: the client's per-call timeout does not cut it off.
Example ¶
package main
import (
"context"
"io"
"log"
"os"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
s, err := vx.Stream(ctx, voixa.StreamRequest{VoiceID: "vi-truc-ly", Text: "Xin chào, Voixa có thể giúp gì?", SampleRate: 16000})
if err != nil {
log.Fatal(err)
}
defer s.Close()
// 16-bit mono PCM at s.SampleRate, clause by clause.
if _, err := io.Copy(os.Stdout, s); err != nil {
log.Fatal(err)
}
}
Output:
func (*Client) Transcribe ¶
func (c *Client) Transcribe(ctx context.Context, req TranscribeRequest) (*Transcript, error)
Transcribe uploads audio and starts a transcription: one API call plus a direct upload of the bytes, which starts the job the moment they land.
It returns as soon as the upload finishes, always with StatusProcessing: the work is as long as the audio (about 12 minutes of machine time per hour of speech, spread over several machines) and a cold queue spends two to three minutes starting a machine first. Poll GetTranscript, or call TranscribeFile, which waits for you.
func (*Client) TranscribeFile ¶
func (c *Client) TranscribeFile(ctx context.Context, req TranscribeRequest, wait ...WaitOptions) (*Transcript, error)
TranscribeFile is Transcribe followed by waiting until the transcript is ready (default: every 5 seconds for up to 2 hours). The result carries Text and Segments. A failed transcript returns an *Error with CodeTranscribeFailed.
Example ¶
package main
import (
"context"
"fmt"
"log"
"os"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
f, err := os.Open("interview.mp3")
if err != nil {
log.Fatal(err)
}
defer f.Close()
t, err := vx.TranscribeFile(ctx, voixa.TranscribeRequest{Audio: f, FileName: "interview.mp3", Title: "Interview"})
if err != nil {
log.Fatal(err)
}
fmt.Println(t.Language, t.DurationSec, t.Text)
fmt.Println(t.SRTURL) // signed for 1 hour
}
Output:
func (*Client) UnshareClip ¶
UnshareClip withdraws the public page; the link stops working.
func (*Client) UpdateVoice ¶
UpdateVoice renames or describes one of your clones. Shared: true makes it usable by every Voixa account, false withdraws it.
func (*Client) Usage ¶
Usage returns the token ledger, newest first: every spend, refund and top-up.
func (*Client) Voices ¶
Voices returns the built-in voices, your own clones and clones other accounts shared. An empty language returns every language.
func (*Client) WaitForClip ¶
func (c *Client) WaitForClip(ctx context.Context, clipID string, wait ...WaitOptions) (*Clip, error)
WaitForClip polls a clip until it is no longer processing (default: every 2 seconds for up to 3 minutes).
func (*Client) WaitForDesign ¶
func (c *Client) WaitForDesign(ctx context.Context, designID string, wait ...WaitOptions) (*VoiceDesign, error)
WaitForDesign polls a design until every option is done (default: every 4 seconds for up to 10 minutes).
func (*Client) WaitForTranscript ¶
func (c *Client) WaitForTranscript(ctx context.Context, transcriptID string, wait ...WaitOptions) (*Transcript, error)
WaitForTranscript polls a transcript until it is no longer processing (default: every 5 seconds for up to 2 hours). DoneSec and DurationSec report progress.
func (*Client) WaitForVoice ¶
func (c *Client) WaitForVoice(ctx context.Context, voiceID string, wait ...WaitOptions) (*Voice, error)
WaitForVoice polls a voice until it is no longer processing (default: every 2 seconds for up to 3 minutes).
type Clip ¶
type Clip struct {
ClipID string `json:"clipId"`
VoiceID string `json:"voiceId"`
VoiceName string `json:"voiceName"`
Language Language `json:"language"`
// Text is the text as read, after the API cleaned it.
Text string `json:"text"`
Chars int `json:"chars"`
// Project is the content group this clip belongs to, if one was given.
Project string `json:"project,omitempty"`
// Source: SourceAPI (API key) or SourceStudio (signed-in web app).
Source Source `json:"source,omitempty"`
// Kind is "podcast" for clips made with CreatePodcast; empty for ordinary clips.
Kind string `json:"kind,omitempty"`
Title string `json:"title,omitempty"`
// Parts of a podcast; PartsDone is reported while the podcast is processing.
Parts int `json:"parts,omitempty"`
PartsDone int `json:"partsDone,omitempty"`
// Engine is always "voixa-1".
Engine string `json:"engine"`
Format string `json:"format"`
SampleRate int `json:"sampleRate,omitempty"`
// DurationSec is the length of the audio in seconds.
DurationSec float64 `json:"durationSec,omitempty"`
// SynthMs is the synthesis time on the engine, milliseconds.
SynthMs float64 `json:"synthMs,omitempty"`
// Tokens charged for this clip. Refunded in full if the clip fails.
Tokens int `json:"tokens,omitempty"`
Status Status `json:"status"`
FailReason string `json:"failReason,omitempty"`
// URL is a signed WAV URL, valid for 1 hour. Empty while processing.
URL string `json:"url,omitempty"`
// ShareURL is the public share page, present after ShareClip. Lives until
// UnshareClip.
CreatedAt string `json:"createdAt"`
UpdatedAt string `json:"updatedAt"`
}
Clip is one generated audio file: a Speak result or a podcast.
type ClipPage ¶
ClipPage is one page of ListClips. Pass Cursor back to read the next page; it is empty on the last page.
type CreateVoiceRequest ¶
type CreateVoiceRequest struct {
Name string `json:"name"`
Language Language `json:"language"`
// Audio is the reference recording: 3 to 10 seconds of clean speech, WAV or MP3,
// at most 10 MB.
Audio io.Reader `json:"-"`
// Size of Audio in bytes. Optional: taken from the reader when it can tell
// (*os.File, *bytes.Reader, *bytes.Buffer, *strings.Reader); otherwise the
// audio is read into memory first.
Size int64 `json:"-"`
// Format is "wav" (default) or "mp3".
Format AudioFormat `json:"-"`
Gender Gender `json:"gender,omitempty"`
Description string `json:"description,omitempty"`
// RefText is what the speaker says in the recording. Improves Chinese, Japanese
// and Korean clones a lot; optional elsewhere.
RefText string `json:"refText,omitempty"`
}
CreateVoiceRequest is the input of CreateVoice.
type DesignCandidate ¶
type DesignCandidate struct {
Index int `json:"index"`
Status Status `json:"status"`
DurationSec float64 `json:"durationSec,omitempty"`
// URL is a signed URL of the option (1 hour). Empty while generating, if it
// failed, or once the design expired.
URL string `json:"url,omitempty"`
}
DesignCandidate is one generated option of a voice design.
type DesignPage ¶
type DesignPage struct {
Items []VoiceDesign `json:"items"`
Cursor string `json:"cursor,omitempty"`
}
DesignPage is one page of ListDesigns.
type DesignVoiceRequest ¶
type DesignVoiceRequest struct {
// Description of the voice in plain words: age, gender, timbre, pace, mood. Any
// language works; English, Vietnamese and Chinese are understood best. Example:
// "An old man in his seventies, hoarse and deep, speaking slowly".
Description string `json:"description"`
// Language of the voice. Must be one Languages reports as Designable.
Language Language `json:"language"`
// Text the options read (max 160 characters, about 8 to 10 seconds; enough to
// clone from). Defaults to the language's sample sentence.
Text string `json:"text,omitempty"`
// Count is how many options to generate, 1 to 3 (default 3). Each costs
// TokenRates.Design tokens.
Count int `json:"count,omitempty"`
}
DesignVoiceRequest is the input of DesignVoice.
type Error ¶
type Error struct {
// Status is the HTTP status (or the status the API would return for a
// client-side check).
Status int
// Message is the human-readable explanation from the API.
Message string
// Code is a machine-readable code (see the Code constants), empty when the
// API did not send one.
Code string
// RetryAfter is how long to wait before calling again, from the Retry-After
// header; zero when absent.
RetryAfter time.Duration
// ResetsAt is when a daily counter resets (ISO 8601 UTC), for legacy quota errors.
ResetsAt string
// Used and Limit accompany legacy quota errors.
Used *int64
Limit *int64
// Needed, Available and RefillsAt accompany CodeInsufficientTokens: the tokens
// the request needs, the tokens you have, and when free tokens refill.
Needed *int64
Available *int64
RefillsAt string
}
Error is returned for every failed API call, for failed uploads and for the client-side checks that mirror the API's validation. Use errors.As to read it:
var ve *voixa.Error
if errors.As(err, &ve) && ve.Status == 404 { … }
type Language ¶
type Language string
Language is a language Voixa speaks. Languages returns labels and clone-recording guides.
type LanguageInfo ¶
type LanguageInfo struct {
Code Language `json:"code"`
Label string `json:"label"`
// MaxChars per Speak call in this language.
MaxChars int `json:"maxChars"`
// Cloneable is false for languages without voice cloning (Chinese, Japanese,
// Korean).
Cloneable bool `json:"cloneable"`
// Designable is true where DesignVoice works.
Designable bool `json:"designable"`
// SampleText is the sentence used for voice samples.
SampleText string `json:"sampleText"`
// Clone is the guide for recording a clone reference.
Clone struct {
Hint string `json:"hint"`
Script string `json:"script"`
} `json:"clone"`
}
LanguageInfo describes one language Voixa speaks.
type ListClipsOptions ¶
type ListClipsOptions struct {
VoiceID string
Project string
Source Source
Kind ClipKind
Limit int
Cursor string
}
ListClipsOptions filter ListClips. Zero values mean no filter.
type ListTranscriptsOptions ¶
type ListTranscriptsOptions struct {
Project string
Source Source
Status Status
Limit int
Cursor string
}
ListTranscriptsOptions filter ListTranscripts. Zero values mean no filter.
type Option ¶
type Option func(*Client)
Option configures a Client.
func WithBaseURL ¶
WithBaseURL points the client at another deployment of the API.
func WithHTTPClient ¶
WithHTTPClient sets the HTTP client used for every request, including uploads and speech streams. Do not give it a short overall Timeout if you use Stream: the timeout would cut long streams off. Use WithTimeout for API calls instead.
func WithIDToken ¶
WithIDToken authenticates as a signed-in Studio user with their ID token (not the access token) instead of an API key. Pass an empty API key to NewClient with it.
func WithTimeout ¶
WithTimeout sets the timeout of each API call (default 30 seconds). Uploads and stream bodies are bounded only by the context you pass.
func WithUserAgent ¶
WithUserAgent appends a product token to the User-Agent header, for example "my-app/1.2".
type PageOptions ¶
PageOptions page through a list. Zero values use the API defaults.
type PodcastRequest ¶
type PodcastRequest struct {
VoiceID string `json:"voiceId"`
// Title is shown in the library and on the public page; defaults to the first
// words of the text.
Title string `json:"title,omitempty"`
// Text, up to 100 000 characters (about 80 minutes). Blank lines separate
// paragraphs (a 600 ms pause); other whitespace collapses. Voixa splits the text
// into sentences, produces the parts in parallel and joins them into one WAV.
Text string `json:"text"`
Project string `json:"project,omitempty"`
}
PodcastRequest is the input of CreatePodcast and MakePodcast.
type Project ¶
type Project struct {
Project string `json:"project"`
Clips int `json:"clips"`
Chars int `json:"chars"`
UpdatedAt string `json:"updatedAt,omitempty"`
}
Project is a content group with its clip and character counts.
type STTInfo ¶
type STTInfo struct {
Languages []struct {
Code string `json:"code"`
Label string `json:"label"`
} `json:"languages"`
Formats []AudioFormat `json:"formats"`
MaxBytes int64 `json:"maxBytes"`
MaxSeconds int64 `json:"maxSeconds"`
Tasks []TranscribeTask `json:"tasks"`
Engine string `json:"engine"`
}
STTInfo lists the languages you may pin on Transcribe, the accepted formats and the limits.
type SampleVoice ¶
type SampleVoice struct {
VoiceID string `json:"voiceId"`
Language Language `json:"language"`
Name string `json:"name"`
Gender Gender `json:"gender"`
Styles []string `json:"styles"`
Kind VoiceKind `json:"kind"`
Description string `json:"description,omitempty"`
// SampleURL is a public URL of a short WAV sample (cached for a year; the file
// never changes).
SampleURL string `json:"sampleUrl"`
// SampleText is the sentence spoken in the sample.
SampleText string `json:"sampleText"`
Engine string `json:"engine"`
}
SampleVoice is a built-in voice with a public sample, for a voice picker in your own app.
type Samples ¶
type Samples struct {
Languages []struct {
Code Language `json:"code"`
Label string `json:"label"`
} `json:"languages"`
SampleText map[Language]string `json:"sampleText"`
Voices []SampleVoice `json:"voices"`
}
Samples is the answer of Samples.
type SaveDesignRequest ¶
type SaveDesignRequest struct {
// Candidate is the 0-based index of the option.
Candidate int `json:"candidate"`
Name string `json:"name"`
Gender Gender `json:"gender,omitempty"`
Description string `json:"description,omitempty"`
}
SaveDesignRequest picks one option of a design to keep as a voice.
type Source ¶
type Source string
Source tells where something was created: with an API key or in Studio (the signed-in web app).
type SpeakRequest ¶
type SpeakRequest struct {
VoiceID string `json:"voiceId"`
// Text, up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Line
// breaks and repeated spaces collapse to one space before counting.
Text string `json:"text"`
// Project is an optional content group, e.g. "episode-12" or "promo Q4". Clips
// with the same project can be listed together and counted (Projects). Letters,
// digits, spaces, ". _ -"; max 60.
Project string `json:"project,omitempty"`
// Wait up to about 20 seconds on the server for the clip. Nil means true; use
// Bool(false) to get the clip back at once, still processing.
Wait *bool `json:"wait,omitempty"`
}
SpeakRequest is the input of Speak and Say.
type SpeechStream ¶
type SpeechStream struct {
io.ReadCloser
// SampleRate of the audio, in Hz.
SampleRate int
// Format of the body: StreamPCM or StreamWAV.
Format StreamFormat
// Tokens charged when the stream opened. Refunded if generation fails.
Tokens int
}
SpeechStream is Vietnamese speech delivered while it is being voiced. It is an io.ReadCloser over the audio: 16-bit little-endian mono PCM (or the same PCM behind a WAV header), clause by clause. Read it until io.EOF and Close it.
type StreamFormat ¶
type StreamFormat string
StreamFormat is the body format of a speech stream.
const ( // StreamPCM is raw 16-bit little-endian mono PCM (the default). StreamPCM StreamFormat = "pcm" // StreamWAV is the same PCM behind a streaming WAV header. StreamWAV StreamFormat = "wav" )
type StreamRequest ¶
type StreamRequest struct {
// VoiceID of a Vietnamese voice (preset or your clone). For other languages use
// Speak.
VoiceID string `json:"voiceId"`
// Text, up to 5 000 characters. Cleaned and counted like Speak.
Text string `json:"text"`
// SampleRate of the output: 48000, 24000 (default), 16000 or 8000. 8000 and
// 16000 suit telephony.
SampleRate int `json:"sampleRate,omitempty"`
// Format: StreamPCM (default) or StreamWAV.
Format StreamFormat `json:"format,omitempty"`
}
StreamRequest is the input of Stream.
type TokenRates ¶
type TokenRates struct {
// ViFast per character, Vietnamese.
ViFast float64 `json:"viFast"`
// Standard per character: English, French, German, Italian, Spanish, Portuguese.
Standard float64 `json:"standard"`
// CJK per character: Chinese, Japanese, Korean.
CJK float64 `json:"cjk"`
// STTPerSec per second of transcribed audio.
STTPerSec float64 `json:"sttPerSec"`
// Clone per voice clone created (also charged when you save a designed voice).
Clone float64 `json:"clone"`
// Design per option generated by DesignVoice (3 options cost 3 times this).
Design float64 `json:"design"`
// Stream per character, Vietnamese streaming.
Stream float64 `json:"stream"`
}
TokenRates say how many tokens one unit of work costs. One token is one Vietnamese character read on the fast path.
func (TokenRates) EstimateClone ¶
func (r TokenRates) EstimateClone() int
EstimateClone is the cost of creating a voice clone or saving a designed voice.
func (TokenRates) EstimateDesign ¶
func (r TokenRates) EstimateDesign(options int) int
EstimateDesign is the cost of a voice design with the given number of options.
func (TokenRates) EstimateSpeak ¶
func (r TokenRates) EstimateSpeak(language Language, chars int) int
EstimateSpeak is the cost of reading chars characters in language, for Speak or CreatePodcast.
Example ¶
package main
import (
"context"
"fmt"
"log"
voixa "github.com/vovix-ai/voixa-go"
)
func main() {
ctx := context.Background()
vx := voixa.NewClient("")
acc, err := vx.Account(ctx)
if err != nil {
log.Fatal(err)
}
if acc.Tokens != nil {
fmt.Println(acc.Tokens.Rates.EstimateSpeak(voixa.Vietnamese, 1200))
}
}
Output:
func (TokenRates) EstimateStream ¶
func (r TokenRates) EstimateStream(chars int) int
EstimateStream is the cost of streaming chars characters.
func (TokenRates) EstimateTranscription ¶
func (r TokenRates) EstimateTranscription(seconds float64) int
EstimateTranscription is the cost of transcribing seconds of audio. Transcription is charged per second of audio after it finishes.
type TokenWallet ¶
type TokenWallet struct {
// Available is spendable right now (free today + balance). Nil for accounts
// that are not limited.
Available *int64 `json:"available"`
// FreeToday is the free tokens left today. They refill to DailyFree at
// RefillsAt; unused free tokens do not carry over.
FreeToday int64 `json:"freeToday"`
DailyFree int64 `json:"dailyFree"`
// Balance is tokens added to the account, spent after today's free tokens. Can be
// slightly negative after a long transcription.
Balance int64 `json:"balance"`
// UsedToday is tokens spent today, net of refunds.
UsedToday int64 `json:"usedToday"`
// RefillsAt is when free tokens refill (ISO 8601 UTC).
RefillsAt string `json:"refillsAt"`
Rates TokenRates `json:"rates"`
Unlimited bool `json:"unlimited"`
}
TokenWallet is a free allowance that refills daily plus a balance that carries over.
type TranscribeRequest ¶
type TranscribeRequest struct {
// Audio is the audio (or video) file.
Audio io.Reader
// Size of Audio in bytes. Optional: taken from the reader when it can tell
// (*os.File, *bytes.Reader, *bytes.Buffer, *strings.Reader); otherwise the
// audio is read into memory first.
Size int64
// Format of the container. Inferred from FileName (or the name of an *os.File)
// when empty; MP3 otherwise.
Format AudioFormat
// FileName is the original file name: shown in the library, not used as a key.
FileName string
Title string
// Language pins the spoken language ("vi", "en", "ja", …). Leave it empty to let
// Voixa detect it, which it does well; see STTLanguages for the list.
Language string
// Task: TaskTranslate recognises the speech and returns English in one pass.
Task TranscribeTask
// Words asks for per-word timestamps (about 10% slower).
Words bool
// Prompt gives spelling hints for names and jargon, read as if it preceded the audio.
Prompt string
Project string
}
TranscribeRequest is the input of Transcribe and TranscribeFile.
type TranscribeTask ¶
type TranscribeTask string
TranscribeTask: TaskTranscribe keeps the spoken language; TaskTranslate recognises the speech and returns English in one pass.
const ( TaskTranscribe TranscribeTask = "transcribe" TaskTranslate TranscribeTask = "translate" )
type Transcript ¶
type Transcript struct {
TranscriptID string `json:"transcriptId"`
Title string `json:"title,omitempty"`
// FileName is the file name you passed, for your own bookkeeping.
FileName string `json:"fileName,omitempty"`
Format AudioFormat `json:"format"`
// RequestedLanguage is the language you asked for, if you pinned one.
RequestedLanguage string `json:"requestedLanguage,omitempty"`
// Language of the speech: detected when you did not pin one.
Language string `json:"language,omitempty"`
// LanguageProb is the confidence of the language detection, 0-1.
LanguageProb float64 `json:"languageProb,omitempty"`
Task TranscribeTask `json:"task"`
Words bool `json:"words"`
Source Source `json:"source,omitempty"`
Project string `json:"project,omitempty"`
// Engine is always "voixa-stt-1".
Engine string `json:"engine"`
AudioBytes int64 `json:"audioBytes"`
// DurationSec is the length of the audio (known once the file is decoded).
DurationSec float64 `json:"durationSec,omitempty"`
// Chars of transcript text.
Chars int `json:"chars,omitempty"`
SegmentCount int `json:"segmentCount,omitempty"`
// TranscribeMs is machine time spent recognising, summed over all parts, so it
// exceeds the time you waited.
TranscribeMs float64 `json:"transcribeMs,omitempty"`
// RTF is machine time divided by audio length. Below 1 is faster than real time
// on one machine.
RTF float64 `json:"rtf,omitempty"`
// Parts: long files are recognised in parallel on several machines and joined
// into one transcript at the end.
Parts int `json:"parts,omitempty"`
// PartsDone so far, while processing.
PartsDone int `json:"partsDone,omitempty"`
// DoneSec is the seconds of audio already recognised, while processing.
DoneSec float64 `json:"doneSec,omitempty"`
// JobStatus is what the job is doing while processing: RUNNABLE (waiting for a
// machine), STARTING, RUNNING. A cold queue takes two to three minutes to reach
// RUNNING.
JobStatus string `json:"jobStatus,omitempty"`
Status Status `json:"status"`
FailReason string `json:"failReason,omitempty"`
// AudioURL is a signed URL of the audio you uploaded, valid for 1 hour.
AudioURL string `json:"audioUrl,omitempty"`
// Signed URLs of the results, valid for 1 hour. Only when you read one
// transcript (GetTranscript), not in lists.
JSONURL string `json:"jsonUrl,omitempty"`
SRTURL string `json:"srtUrl,omitempty"`
VTTURL string `json:"vttUrl,omitempty"`
TXTURL string `json:"txtUrl,omitempty"`
// Text is the full text. Only when you read one transcript, not in lists.
Text string `json:"text,omitempty"`
// Segments are the timed segments. Only when you read one transcript.
Segments []TranscriptSegment `json:"segments,omitempty"`
CreatedAt string `json:"createdAt"`
UpdatedAt string `json:"updatedAt"`
}
Transcript is the result of a transcription.
type TranscriptPage ¶
type TranscriptPage struct {
Items []Transcript `json:"items"`
Cursor string `json:"cursor,omitempty"`
}
TranscriptPage is one page of ListTranscripts.
type TranscriptSegment ¶
type TranscriptSegment struct {
ID int `json:"id"`
// Start and End are seconds from the start of the audio.
Start float64 `json:"start"`
End float64 `json:"end"`
Text string `json:"text"`
// Words is set only when the request asked for Words.
Words []TranscriptWord `json:"words,omitempty"`
}
TranscriptSegment is a timed piece of the transcript.
type TranscriptWord ¶
type TranscriptWord struct {
Start float64 `json:"start"`
End float64 `json:"end"`
Word string `json:"word"`
// Prob is the confidence, 0-1.
Prob float64 `json:"prob"`
}
TranscriptWord is one word with its timing, when the request asked for Words.
type UsageEntry ¶
type UsageEntry struct {
ID string `json:"id"`
// Kind: speak, podcast, clone, design, transcribe, stream, refund or grant.
Kind string `json:"kind"`
Tokens int64 `json:"tokens"`
// Free and Paid split Tokens between today's free tokens and the balance.
Free int64 `json:"free"`
Paid int64 `json:"paid"`
// Ref is the clipId, voiceId or transcriptId concerned.
Ref string `json:"ref,omitempty"`
Note string `json:"note,omitempty"`
At string `json:"at"`
}
UsageEntry is one line of the token ledger: negative for spends, positive for refunds and top-ups.
type UsagePage ¶
type UsagePage struct {
Items []UsageEntry `json:"items"`
Cursor string `json:"cursor,omitempty"`
}
UsagePage is one page of Usage.
type Voice ¶
type Voice struct {
VoiceID string `json:"voiceId"`
// Kind is VoiceKindPreset for a built-in voice, VoiceKindClone for one created
// from a reference recording (yours, or a built-in one).
Kind VoiceKind `json:"kind"`
// Mine is true for voices your account created: those can be renamed, shared
// and deleted.
Mine bool `json:"mine"`
// Shared: a clone its owner shared with every account (or your own clone, if
// you shared it).
Language Language `json:"language"`
Name string `json:"name"`
Gender Gender `json:"gender"`
Styles []string `json:"styles"`
Description string `json:"description,omitempty"`
// SampleURL is a short sample of the voice. Signed URL for clones (1 hour);
// empty while a clone is still processing.
SampleURL string `json:"sampleUrl,omitempty"`
Status Status `json:"status"`
FailReason string `json:"failReason,omitempty"`
// Engine is always "voixa-1".
Engine string `json:"engine"`
// DesignID is set when the voice was saved from a voice design (SaveDesign).
DesignID string `json:"designId,omitempty"`
CreatedAt string `json:"createdAt"`
UpdatedAt string `json:"updatedAt"`
}
Voice is a built-in voice or a clone.
type VoiceDesign ¶
type VoiceDesign struct {
DesignID string `json:"designId"`
Language Language `json:"language"`
Description string `json:"description"`
// Instruct is what the voice engine was given: your description, translated to
// English when needed.
Instruct string `json:"instruct,omitempty"`
Text string `json:"text"`
Count int `json:"count"`
Status Status `json:"status"`
FailReason string `json:"failReason,omitempty"`
Candidates []DesignCandidate `json:"candidates"`
// SavedVoiceIDs are the voices saved from this design.
SavedVoiceIDs []string `json:"savedVoiceIds"`
Source Source `json:"source,omitempty"`
Tokens int `json:"tokens,omitempty"`
ExpiresAt string `json:"expiresAt"`
Expired bool `json:"expired"`
CreatedAt string `json:"createdAt"`
UpdatedAt string `json:"updatedAt"`
}
VoiceDesign holds options generated from a description. Options are drafts kept for 23 hours (ExpiresAt); save the one you like with SaveDesign to get a permanent voice.
type VoiceFromDescriptionRequest ¶
type VoiceFromDescriptionRequest struct {
Description string
Language Language
Text string
Name string
Gender Gender
}
VoiceFromDescriptionRequest is the input of CreateVoiceFromDescription.
type VoiceKind ¶
type VoiceKind string
VoiceKind: a built-in preset, or a clone made from a reference recording (yours, or a built-in one).
type VoiceUpdate ¶
type VoiceUpdate struct {
Name string
// Description: a pointer to "" clears it.
Description *string
Gender Gender
// Styles: a non-nil empty slice clears them.
Styles []string
Shared *bool
}
VoiceUpdate changes a clone. Nil or empty fields are left as they are.
func (VoiceUpdate) MarshalJSON ¶
func (u VoiceUpdate) MarshalJSON() ([]byte, error)
MarshalJSON sends only the fields that change.
type WaitOptions ¶
type WaitOptions struct {
// PollInterval between two status reads.
PollInterval time.Duration
// Timeout: give up after this long with an *Error of code CodeTimeout. Cold
// engines can take about a minute.
Timeout time.Duration
}
WaitOptions tune the Wait helpers and the methods that wait for you. Zero fields keep the method's default.