Factor
A fast, lightweight AI agent that lives on your machine β with hands on your
desktop, a voice on the phone, and a real memory.
Factor is a single static Go binary. Talk to it in the terminal, on Telegram, over
the phone, or out loud to the machine itself β same agent, same tools, same memory
every time. It drives a real browser, works your desktop, runs long tasks in the
background, and calls or texts you when one lands. Long-term memory is
smrti: it consolidates and prioritizes rather
than logging, so past failures become constraints it doesn't repeat.
Highlights
|
|
| π§ Memory as the soul |
Salience-ranked recall every turn, consolidation that decays and prunes, plus remember / recall / forget / reflect tools |
| β‘ Never keeps you waiting |
Says what it's about to do while the tools run; long work becomes a background job that messages you when it lands |
| π― Mid-turn steering |
A second message during a live turn is injected between tool iterations, not queued |
| β Asks when only you know |
ask_user puts the question where you are: the chat the turn came from, the terminal you're already in, or a dialog on your desktop β and times out rather than hanging the turn when you're away |
| π Provider failover |
OpenAI-compatible (OpenRouter, Ollama, LM Studio, Groq, llama.cpp, β¦) and native Anthropic, with per-candidate cooldowns and overflow compaction |
| πΈ Counts what it spends |
Every call priced and billed to its session β status bar, tray, usage tool β with per-session and global caps; cache reads and writes priced at their own rates, not as fresh input |
| β»οΈ Caches the prompt prefix |
The same assembly order every turn, with cache breakpoints where two turns first differ, so the provider reuses the longest prefix it can β and a turn twenty tool calls deep stops reprocessing its own history |
| π§ Reasoning, dialect-translated |
One provider.reasoning setting becomes reasoning, reasoning_effort, or a thinking budget |
| βοΈ Answers the phone |
A real number to call, or it calls and texts you β barge-in, voicemail detection, optional fully local speech (Phone) |
| ποΈ Listens in the room |
Mic in, speakers out, barge-in, optional wake word, push-to-talk via factor talk (PC voice) |
| π₯ Tells voices apart |
A recording is read voice by voice, so each person holds their own conversation β and company in the room moves the answer to a memory space it can hear (PC voice) |
| ποΈ Hands on your desktop |
Windows, screenshots, mouse, keyboard, clipboard, notifications on X11/Wayland/macOS/Windows β plus grid vision (Desktop) |
| π A real browser, not just fetch |
CDP tools attach to your running Chrome/Chromium/Brave or launch a managed one (Browser) |
| π§© Extensible everything |
Channel connectors, Go tools, runtime-mounted MCP servers, markdown skills (Extending) |
| π Watches its own numbers |
Every turn leaves a local trace β models, tools, timings, cache and cost β and control bands measure it against a rolling baseline, so the heartbeat wakes a model only once a number has drifted |
| π Learns from its own work |
A turn that took four or more tool calls, or one you had to steer, becomes a skill it writes for itself once the session goes quiet (Extending) |
| π§ Self-managing |
Edits its own config, installs packages, upgrades and restarts itself, schedules cron jobs and one-off reminders, runs HEARTBEAT.md checks that cost nothing when idle |
| π‘οΈ Safety rails |
Workspace-restricted files, exec deny-patterns, sender allowlists, scrubbed secrets β rails, not a sandbox (Security) |
How it works
Bus + bounded workers, mid-turn steering, narrow pluggable seams, CGO-free
portability β PicoClaw's architecture in a
codebase that runs happily on an old Puppy Linux box.
Get started
go install github.com/cyqlelabs/factor/cmd/factor@latest
# or grab a release binary; linux-amd64 targets GOAMD64=v1 (no SSE4.2 needed)
factor init # interactive setup wizard
The wizard verifies every step live β a real provider completion, the endpoint's
model list, Telegram's getMe, carrier and voice credentials, an actual page load β
and installs what's missing: smrti, a browser, your desktop backend's helpers. It
probes the machine's display rather than this shell, so setup over ssh targets the
right desktop, and it can add a login entry (systemd or XDG autostart, launchd, the
Windows Run key). factor init -y takes the defaults; --no-install installs
nothing.
export FACTOR_PROVIDER_API_KEY=sk-or-... # OpenRouter by default
factor # interactive chat
factor -m "what's on my disk?" # one-shot
factor gateway # daemon: Telegram, phone, PC voice, cron, heartbeat, jobs
factor gateway -d # the same, detached (~/.factor/gateway.log)
factor talk # push-to-talk: arm the PC voice microphone
factor status # daemon / provider / memory / phone / voice / desktop health
factor upgrade # replace this binary with the newest release
factor -p 127.0.0.1:8080 # route HTTP through a proxy and watch every call
-p routes Factor's HTTP through any proxy β mitmproxy, Burp, ZAP, SOCKS5 β so you
can read the prompts, tool schemas, replies and token counts it actually sends. On a
desktop, the running gateway also puts a status icon in the system tray: version,
uptime, memory health, connected channels, and a clean quit.
Proxy details, the memory sidecar, and how upgrades land
Loopback stays direct and child processes inherit the proxy setting, so smrti's calls
show up but the local sidecars aren't caught. --proxy-ca trusts an intercepting
proxy's CA, probed once at startup. A flag only reaches the process you typed it at,
and the gateway that systemd or your login starts gets none, so proxy.address and
proxy.ca in the config (FACTOR_PROXY, FACTOR_PROXY_CA) carry the setting to
every start; a typed -p still wins. An address is tried with one request before it is
saved or applied on reload, so a proxy nothing answers at is refused instead of
failing every call, and the agent cannot move it from a heartbeat. The browser isn't
routed; it has its own trust store. The tray is absent on a headless box and on
macOS, whose tray would cost the build its CGO-free binaries.
Factor supervises the smrti sidecar, restarts it with backoff, and degrades
gracefully (empty recalls, dropped writes) when it's down. Point
memory.mode: "external" + memory.url at a shared smrti if you run one.
factor upgrade downloads the release for this machine, verifies it against the
published SHA256SUMS, and swaps the binary in place (--check only reports). A
running gateway restarts into it once the turn in flight is answered, keeping its
pid so systemd never sees it stop. Factor checks daily and tells you, never
installing unasked.
The same command brings smrti up to date however it runs here: a container is
recreated on the newly published image, and a uv, pipx, pip or venv install is
upgraded by the installer that made it, then restarted into. Both wait for the
memory graph to go quiet first, so nothing in flight is lost. An engine on
another machine is left to whoever runs it.
Configuration
~/.factor/config.json β every key optional, defaults work. FACTOR_* env
overrides: FACTOR_PROVIDER_API_KEY, FACTOR_PROVIDER_MODEL, FACTOR_MEMORY_MODE, β¦
A running gateway watches the file. Save an edit, by hand or through the agent's own
config_set, and it reloads within seconds β after the turn in flight is answered,
and without restarting the sidecars β then names the changed sections in the chat it
reports back to. A save that doesn't parse is warned about and retried, never applied.
Annotated example
{
"log_level": "info", // debug | info | warn | error
"agent": {
"context_window_tokens": 0, // 0 = ask the model catalog; a value only ever shrinks its answer
"max_tool_iterations": 20, // per stretch; a turn still mid-task is checkpointed and gets up to three
"summarize_at_percent": 75, // how full the window gets before compaction
"keep_recent_messages": 8, // what survives it
"learn_skills": true, // distill a finished multi-tool turn into a skill
"version_workspace": false // keep a local git history of the workspace, so an edit can be undone
},
"provider": {
"type": "openrouter", // openrouter|openai|groq|ollama|lmstudio|llamacpp|anthropic|custom
"api_key": "sk-or-...",
"model": "google/gemini-3.1-pro-preview",
"reasoning": { "effort": "xhigh" }, // or {"max_tokens": 12000}; "none" turns it off
"fallbacks": [{ "type": "ollama", "model": "qwen3:8b" }],
"utility": [{ "type": "ollama", "model": "qwen3:8b" }] // cheaper chain for compaction summaries and skill verdicts; omit = the main one
},
"memory": {
"max_rss_mb": 1536, // restart the engine for size once idle; -1 turns it off
"mode": "sidecar", // sidecar | external | off
"auto_install": true, // install smrti when it is missing
"personality": "balanced", // analytical | curious | empathetic | maverick | deterministic
"space": "main", // where conversations are remembered
"space_strategy": "origin", // origin: cron and job turns use system_space | single: one space for all
"system_space": "system",
"shared_space": "shared" // what a turn other people can hear reads and writes
},
"channels": {
"telegram": { "token": "123:ABC", "allow_from": ["your-telegram-id"] },
"phone": { // optional; absent = nothing runs
"user_number": "+15550001111", // you: the only one who may call in
"phone_number": "+15550002222", // the number you bought
"carrier": "twilio", // twilio | telnyx
"twilio_account_sid": "AC...", // twilio: these two
"twilio_auth_token": "...",
// telnyx instead: "telnyx_api_key", "telnyx_connection_id", "telnyx_public_key"
"elevenlabs_api_key": "...",
"stt_api_key": "...", // Deepgram
"language": "en",
"stt": { "provider": "deepgram" }, // deepgram | whisper | local-openai
"tts": { "provider": "elevenlabs" }, // elevenlabs | local-openai
"proactive": "sms", // sms | call | off
"max_call_minutes": 15
},
"voice": { // PC voice: this machine's mic and speakers
"activation": "wake-word", // always | wake-word | push-to-talk
"wake_word": "factor",
"language": "en",
"stt": { "provider": "deepgram" }, // deepgram | whisper | local-openai
"stt_api_key": "...", // Deepgram
"tts": { "provider": "elevenlabs" }, // elevenlabs | local-openai
"elevenlabs_api_key": "...",
"speaker_id": false, // tell the room apart; needs a local speech tier
"speaker_threshold": 0.35, // similarity below which a voice is nobody enrolled
"unknown_speaker": "anonymous", // anonymous | enroll
"room_isolation": null, // null = on wherever speaker_id is
"room_timeout_minutes": 30, // how long a voice counts as still in the room
"output_volume": 100, // 1β100; lower it when the speakers reach the mic
"ignored_chime": true // a soft tone when it heard you and did not take it
}
},
"mcp": {
"servers": { "github": { "command": "github-mcp-server", "args": ["stdio"] } }
},
"tools": { "disabled": [], "restrict_to_workspace": true },
"desktop": { "enabled": null }, // null = on when a display exists
"browser": {
"enabled": true,
"engine": "auto", // auto | camofox | chromium
"command": "", // "" = find one; init records what it installed
"headless": false,
"camofox": { "port": 9377 } // the headless engine's sidecar; one already answering here is adopted
},
"heartbeat": { "enabled": true, "interval_minutes": 30 },
"cron": { "job_timeout_minutes": 30 }, // how long one scheduled task may run; it is told what is left at each checkpoint
"trace": {
"enabled": true, // one JSON line per turn in ~/.factor/traces
"record_args": false, // the shape of a turn, not what was said to the tools
"keep_days": 14
},
"upgrade": { "check": true, "check_interval_hours": 24 }, // report new releases; never install one unasked
"proxy": { "address": "", "ca": "" }, // "" = direct; host:port or a URL routes every call, sidecars included
"cost": {
"track": true, // price every call; models served locally cost nothing
"budget": {
"session_usd": 0, // 0 = no cap, on both scopes
"global_usd": 0,
"period": "month" // what "global" counts: day | month | total
},
"prices": {} // USD per million tokens, for models the catalog does not list
}
}
Spend is priced from the model catalog cached in ~/.factor/pricing.json and
totalled in ~/.factor/usage.json, per session and overall. Locally served models
are free; models the catalog doesn't list are counted in tokens rather than guessed
at. Caps are checked before the call, and the turn answers with a line saying what
stopped. Ask for usage to see the breakdown.
The workspace (~/.factor/workspace) is the agent's home. The persona is built into
the binary, so an upgrade improves it everywhere at once; SOUL.md layers yours on
top, USER.md holds what Factor should always know about you, AGENT.md tunes how
it works, HEARTBEAT.md lists proactive tasks, and instructions/, skills/,
sessions/, cron/ do what they say.
Browser
Two engines answer the same eleven tools β browser_navigate Β· _read Β· _scroll
Β· _click Β· _fill Β· _keys Β· _upload Β· _tabs Β· _screenshot Β· _eval Β·
_back β and browser.engine picks between them:
| Engine |
What it is |
Runs when |
| Chromium β Helium, or the browser you already run |
a real browser driven over DevTools, visible on the desktop |
a display exists, or a browser of yours is attached on port 9222 |
| Camofox β camofox-browser over Camoufox |
a Firefox build whose fingerprint is spoofed in C++ before any page script runs, so Cloudflare and the sites that turn a headless Chromium away serve it |
there is no display to open a window on, or browser.headless is set |
auto is that table; camofox and chromium force one engine. Browse as
yourself. Start your everyday browser with --remote-debugging-port=9222 (or point
browser.attach_url at it) and Factor uses that session instead of launching one β
your logins, your cart, your cookies β and it wins over both engines.
Nothing here is installed by hand. factor init puts both engines down β Helium for a
machine with no browser, Camofox on every machine β and a gateway that finds Camofox
missing installs it in the background as it starts, so an install upgraded from a Factor
without it has the engine before the first page is asked for. The Camofox install is an
npm package and a 300 MB Firefox build kept under ~/.factor/engine/camofox, run on the
machine's Node 22.13 or newer, or on a Node Factor downloads for Linux, macOS or Windows
and checks against its published checksum. Two of the package's native bindings want a
glibc newer than the distributions Factor is most often put on; where they will not
load, Factor stands pure-JavaScript stand-ins in for them (the browser itself needs
nothing newer than glibc 2.18), so no compiler is ever needed. It runs as a sidecar on
port 9377 β one you run yourself is adopted β and idles at a sleeping Node process,
launching Firefox on the first page and shutting it down when nothing has asked for a
while. Its crash telemetry is switched off.
browser_read says how much it withheld and takes filter/limit; on Camofox a page
comes back as an accessibility snapshot with element refs, a tenth the size of the HTML,
and offset reads on past a cut. browser_scroll reaches what only loads on the way
down.
Desktop
Factor works the graphical session through the desktop's own helper programs
(xdotool, wmctrl, scrot on X11; grim/wtype on Wayland; osascript on macOS;
PowerShell on Windows) β no CGO bindings, and nothing registers on a headless box.
On top of the window/mouse/keyboard/clipboard tools sits grid vision, a two-pass
pointing loop for vision-capable models:
screen_view captures the screen under a battleship coordinate grid β columns
A, B, Cβ¦, rows 1, 2, 3β¦. The model names the cell it sees ("the icon is in D4")
instead of guessing pixel coordinates, which vision models are bad at.
screen_zoom cell=D4 magnifies that cell β or any pixel region, e.g. a window's
geometry from window_list β under a finer sub-grid, down to ~10px precision.
mouse action=click cell=B3 clicks the cell's center, resolved back to native
screen pixels on either view.
Pure Go image math β no OpenCV, no OCR, no helpers beyond the screenshot program.
Frames are capped at 1568px on the longest side (clicks still land at native
resolution), only the two newest stay in context, and image bytes never touch
session history. Non-vision models can disable both vision tools via
tools.disabled.
Phone calls and SMS
Give Factor a phone number and it picks up: you talk, it answers out loud, with the
same memory, tools, and history it has everywhere else. It can also call or text
you β a finished job, a cron result, or because you asked it to ring someone.
factor init # the Channels step walks through the carrier and the speech tier
factor gateway # brings the line up
factor status # number, speech tier, voice-shell health
Carrier setup has a page of its own: Twilio Β·
Telnyx.
The voice shell is Patter, installed on demand
and supervised like the smrti sidecar. It owns the carrier's media stream,
turn-taking, barge-in, voice activity detection, answering-machine detection and
transcoding; Factor is its brain, over an endpoint that never leaves 127.0.0.1.
Speech tiers β the one decision with real trade-offs, asked by the wizard:
| Tier |
Speech-to-text |
Text-to-speech |
Extra RAM |
When |
| 1 Β· cloud (default) |
Deepgram nova-3 |
ElevenLabs flash v2.5 (Β΅-law 8 kHz, no transcode) |
~150β300 MB |
any machine; lowest latency, least to go wrong |
| 2 Β· local STT |
Parakeet / Whisper |
ElevenLabs |
+0.5β2 GB |
transcription stays home |
| 3 Β· local TTS |
Deepgram |
Piper |
+0.3 GB |
Piper's ~100 ms render beats the cloud, on any CPU |
| 4 Β· fully local |
Parakeet / Whisper |
Piper |
+1β2 GB |
no audio leaves the machine, no per-minute audio cost |
A tier picks who serves each half of the pipeline; everything else is the same call:
Pick a local tier and Factor installs it: engines in their own virtualenv, your
language's models on disk before setup finishes, nothing to start by hand.
Roughly $0.04β0.06 per talk-minute on tier 1 plus your model's tokens, and about
1.3Β’ per SMS segment. Simple questions land in 1.5β3 s. The reply arrives whole, but
a tool-using turn still speaks: the line Factor says on its way to the answer is
streamed into the live call while the tools run.
Which languages and voices you get, and which transcriber
Transcription covers Whisper's ~99 languages, with Parakeet serving its 25 at higher
accuracy where the machine allows. Voices come from Piper's catalogue of 49,
resolved from your language setting with the exact locale winning where it exists
(es-MX gets a Mexican voice, not a Castilian one). You can also pick the voice by
name: the wizard lists your ElevenLabs voices on the cloud tier and the catalogue's on
the local one, and a voice named in speech_server.piper_voice is downloaded on the
next start.
speech_server.speech_speed paces it: 0.9 speaks a tenth slower, 1.1 a tenth faster.
Piper stretches the phonemes rather than the audio, so the voice keeps its pitch β
past about 1.2 either way it stops sounding like a person.
Whisper decodes a fixed 30-second window however little audio it gets, and the phone
pipeline feeds it about a second at a time β so its cost is per chunk, not per second
of speech. Measured on this design: small takes ~2.4 s per 1 s chunk on a CPU and
falls behind while you talk; base keeps up at ~0.9 s but mishears more. That used
to be the CPU's ceiling.
Parakeet TDT 0.6B v3 breaks the
trade: a transducer's cost scales with the audio it is handed, not a fixed window,
and its accuracy benchmarks at Whisper large-v3 level. So the installer picks:
| Machine |
Transcriber |
| GPU |
Whisper large-v3-turbo β every language, large-class accuracy |
| CPU, β₯4 GB RAM, one of Parakeet's 25 languages |
Parakeet TDT int8 (~1 GB resident) |
| CPU otherwise |
Whisper base/tiny, as before |
speech_server.stt_engine (parakeet / whisper) or speech_server.whisper_model
override the choice. On a machine that lands on base, tier 3 β Piper's ~100 ms
render locally, transcription in the cloud β is still the better trade, and the
wizard says so.
Bring your own speech server
Point stt.base_url / tts.base_url at anything OpenAI-compatible β
Speaches or your own β and Factor uses it
as-is. It probes at startup either way and falls back to the cloud tier if the
server isn't answering; set local_audio_fallback: false to have the channel report
itself down instead. Silero voice-activity detection runs locally in every tier.
Getting the line up, and the guardrails on it
Buy a number at Twilio or Telnyx, then run factor init. The wizard asks which one,
takes the credentials, and verifies them live before writing anything.
|
Twilio |
Telnyx |
| Credentials |
account SID + auth token |
API key + connection id + public key |
| Setup at the carrier |
buy a number |
buy a number, create a Call Control Application |
| Cost |
the baseline above |
lower per minute and per text |
| Step by step |
docs/phone-twilio.md |
docs/phone-telnyx.md |
The carrier is pointed at the shell on every start β Patter does it for Twilio,
Factor for Telnyx β so a rotating tunnel keeps working with nothing to click.
The rails are closed by default, because a number is dialable by anyone and every
minute costs money: only user_number may call in, only user_number may be
dialed, calls are cut off at max_call_minutes, and transfer is off. Two tools
appear once the channel is configured β phone_sms sends a text and phone_call
dials, reporting the outcome with a transcript tail into the conversation that asked.
Your carrier's page has the rest: the allowlist knobs that widen those rails, how to
move off the default Cloudflare quick tunnel β fine for a first call, wrong for
daily use β and what to check when the line does not come up.
PC voice: mic and speakers
The same conversation without a phone bill: Factor listens on the machine's own
microphone and answers through its speakers. The mic opens whenever Factor runs β in
factor (the terminal chat keeps working alongside) and in factor gateway alike.
factor init # the Channels step sets up the mic, the speech tier, and the activation
factor # or factor gateway β either one listens
factor talk # push-to-talk: arm the microphone from any terminal
factor status # tier, activation, helpers, and whether anything is listening
activation |
It answers |
always |
every utterance β best alone in a quiet room |
wake-word |
utterances that open with the wake word, plus a short window after each reply so follow-ups don't need it (the wizard preselects this) |
push-to-talk |
nothing until factor talk arms the microphone |
factor talk works in every mode: it rescues a misfired wake word and cuts off
whatever is playing, and /talk does the same inside the chat. Talk over a reply and
it stops mid-word, cancelling the turn behind it. The speech tiers are the phone's,
chosen the same way, with the local server on its own port so both channels can keep
speech at home.
It hears people, not clips. Two people talking to each other leave gaps far
shorter than the silence that closes a recording, so their whole exchange arrives as
one clip β and a single embedding of that is a blend that hands both of them one
name. With speaker_id on, the speech server separates the voices first and reads
each one alone. The owner keeps the main conversation, a recognized guest gets a
session of their own and is named to the agent as the person speaking, and
voice_speakers turns a profile born speaker-2 into Roxana.
And it knows who is listening. That, not who asked, is the question
confidentiality turns on. A second voice in a recording makes the room shared β at
most one of them is the owner, so the rest are company β which holds before anyone is
enrolled and even when nobody could be named, so long as the voice holds a second of
speech of its own.
| Room |
Session |
Recalls from |
Remembers into |
| private |
voice:local, or the guest's own |
space + shared_space |
space |
| shared |
voice:local:room β everyone in one thread |
shared_space |
shared_space |
Three seconds of a voice that is not yours, added up over a couple of minutes,
declares company β a cough, a shout or the "Gracias." Whisper hears in noise never
gets there, and a guest does within their first few remarks; room_timeout_minutes
of silence or the room tool takes it back. The asymmetry is deliberate: a room
wrongly called shared costs you a coy answer, one wrongly called private says
something private to a guest.
Microphone, meter and barge-in
Audio rides the sound system's own helpers β parec/paplay, pw-record/pw-play,
arecord/aplay, sox's rec/play on macOS, sox's waveaudio driver on Windows β
installed by the wizard, which also asks which microphone and proves it live: you
make a noise, it measures, and a silent source is called out on the spot. The chat's
status bar carries a live meter β mic βββ moves with the room and turns green on
speech, βͺ lights cyan while Factor talks, a dead source shows mic β.
Voice activity detection is pure Go: adaptive noise floor, a pre-roll so the first
syllable survives, and a higher bar while the agent speaks so the speakers can't
barge in on themselves. Speakers loud enough get past that bar anyway, so what got
through is matched against the words Factor just sent them: its own voice is dropped
instead of answered, or stripped off the front of what you actually said.
output_volume turns the reply down in rooms where the speakers overpower the
microphone.
A sentence the wake word never reached gets a soft two-note chime, so a misfire sounds
different from a machine that never heard you. It is a tone and not a voice, so talking
across it is not an interruption and still needs the wake word; it waits fifteen seconds
between chimes, because a conversation held in front of the machine is turned away
sentence by sentence. ignored_chime turns it off.
How a voice becomes a name
Each voice is matched against the profiles in ~/.factor/voice-speakers.json, which
needs a local speech tier β that is what computes the embeddings. The turn belongs to
whoever opened the recording, since the wake word was at its front; anyone who joined
halfway through is the room, not the asker. unknown_speaker decides what a new voice
gets: a profile on the spot, or the main conversation, unnamed.
Three bars rise with the stakes: one second of speech to put a name to a voice, two
to fold it into that profile, three to create one, because a spurious profile never
goes away and competes for every match after it. Every decision lands in the log with
the similarity behind it, so a turn answered as the wrong person is readable rather
than a guess.
The room is read from every utterance the mic resolves, including the ones the wake
word turned away β someone talking to you is still someone in the room. Sound only
reports people who make it, so the room tool is how you mention someone who came in
quietly or left. Either flip is spoken before the answer that depends on it, and the
room outlives the process, aged against the wall clock rather than uptime, so a 9pm
upgrade doesn't bring Factor back up private while the guest is still on the sofa.
Replies you can listen to
A spoken turn is told it's being heard rather than read, so replies come out sayable
β no markdown, no bullet lists, no spelled-out URLs. voice_write sends anything
long or written to your terminal instead, or to the chat you last used when Factor
runs as a daemon.
The local tier keeps to itself. Factor sets ORT_DISABLE_TELEMETRY=1 in the speech
process's environment before it starts, because onnxruntime otherwise uploads your
OS build, CPU, memory and a persistent device id as it initializes, and its own
disable_telemetry_events() runs too late to stop it
(onnxruntime#25573).
Extending Factor
| Seam |
What it takes |
| Connector |
One package: channel.Register(name, factory) in init(), with its own config section |
| Tool |
Four methods β Name, Description, Parameters, Execute β and one registry.Register(t) line |
| MCP server |
mcp_add (or the mcp.servers config section) mounts its tools at runtime β no Go required |
| Skill |
Drop workspace/skills/<name>/SKILL.md β catalog in prompt, full text on demand, skill_find searches the public registry (skills.sh) and skill_install takes its slug, a git URL, or a directory |
It also writes its own. Factor remembers a turn that took four or more tool
calls, or two if you steered it while it ran, since a correction names both the
approach that was wrong and the one that worked; once that session sits quiet for
ten minutes, it spends one metered call asking whether the trajectory holds a
workflow worth keeping. Most of the time
the answer is SKIP. A LEARN lands as a skill in the same catalog, marked
learned: true, and the next turn lists it like any other. Induction rewrites
its own output and never a skill you wrote or installed β skill_write drops the
marker, which is how a learned skill graduates out of its reach. The learned
library holds 40, so past that induction has to improve an entry rather than
mint a near-duplicate. agent.learn_skills: false, or FACTOR_LEARN_SKILLS=0,
turns it off.
Wiring in a connector
func init() {
channel.Register("mychat", func(raw json.RawMessage, b *bus.MessageBus) (channel.Channel, error) {
var cfg MyConfig
_ = json.Unmarshal(raw, &cfg) // your own config section
return New(cfg, b), nil // implement Name/Start/Stop/Send/MaxMessageLength
})
}
Security model
Factor is a personal agent, not a multi-tenant service. The guardrails (workspace
restriction, exec deny-patterns, allowlists, secret redaction) protect against
accidents and casual prompt-injection β they are not a security boundary. Run it
under your own account for yourself; set channels.telegram.allow_from; keep
restrict_to_workspace on unless you know why you're turning it off. The phone
channel is the one place the default is closed rather than open β a number anyone
can dial, with a bill attached to every minute, earns stricter rails.
Development
make check # gofmt + vet + race tests + coverage gate (β₯90%, what CI runs)
make lint # golangci-lint (its own CI job, so check alone is not the gate)
make hooks # point git at .githooks: every commit lints first
make build # local binary
make build-all # release cross-compile (incl. GOAMD64=v1 for old x86-64)
make build-tiny # -tags nobrowser: smallest binary
make diagrams # re-render docs/assets/*.mmd into the PNGs the docs embed (needs node)
The suite runs against fakes β scripted providers, a fake smrti sidecar and a fake
voice shell (both spawned by re-execing the test binary), a fake Telegram API, a
fake carrier, a scripted microphone and speaker, a fake MCP server over real stdio
JSON-RPC, a scripted desktop β plus live headless-Chrome and desktop round-trip
tests that auto-skip where the machine can't host them.
License
MIT Β© CyqleLabs