Documentation
¶
Overview ¶
Package api — AI observability endpoints (v0.5.163). The /ai page is the Coremetry-native counterpart to Langfuse: every Copilot Explain call lands as a row in the ai_calls CH table and these endpoints surface it as KPIs / timeseries / a recent- calls table. No external service involvement: prompts + samples stay inside the customer's CH cluster.
Package api handler for /admin/alert-tuning — surfaces the noisiest rules in a window plus per-rule actionable suggestions for the v0.5.127-129 dampening knobs. Admin-only read endpoint; cached 5 min because rule-noise pattern doesn't shift minute- to-minute and operators commonly load this page in bursts.
v0.9.613 — /admin/clickhouse DDL kuyruğu sağlık ucu.
Üç gecedir süren prod vakasının ürünleşmiş teşhisi: operatör artık elle clusterAllReplicas sorgusu yazmıyor, panel verdict + eylem cümlesini veriyor. Sorgu ayrıntıları ve ayırıcının kalibrasyonu chstore/ddl_queue_health.go başlığında.
v0.9.655 — korelasyon kimliğinden DIŞ log sistemine köprü.
Operatör: "Request Id bulduğunda prod'ta log izleme linkini de versin parametrik olarak." Ardından: "test ortamlarının log adresleri ayrı … service isimlerinin sonunda -int -prep -uat gördüğünde o adreslere yönlendirsin."
v0.9.580 zaten örnek request_id'leri buluyor ve cevaba yazıyor ("Örnek request_id: SPE0250…"). Ama operatör o değeri KOPYALAYIP kurumun log arayüzüne elle yapıştırıyor. Cevap bir başlangıç noktası veriyordu; tık verilebilirdi.
ŞABLONLAR AYARDIR, KODA GÖMÜLMEZ. İki sebep:
- Her kurulumun log sistemi başka; adres kod tabanına ait değil.
- Bu depo bir müşteri adresi taşımaz.
ORTAM BAŞINA ŞABLON: aynı kurulum birden çok ortamın verisini taşıyabiliyor (project-env-separation) ve her ortamın log arayüzü ayrı bir adres. Ortam SERVİS ADININ SONEKİNDEN çözülüyor — operatörün tarif ettiği gerçek: "-int", "-uat", "-prep".
v0.9.563 — dashboard güncellemesinde alan taşıma.
Bulunuşu: preset dashboard paketini yenilemek için yapılan çekişmeli incelemede, yeni paketin $service değişkenine güvenemeyeceği ortaya çıktı — çünkü değişken KAYDETMEDE siliniyordu.
Hata zinciri (dördü de doğrulandı):
- Frontend kaydederken gövdeye yalnız {name, description, panels} koyuyor — `variables` YOK.
- Handler gövdeyi TAZE bir chstore.Dashboard struct'ına decode ediyor; ileri taşıdığı tek alan CreatedAt.
- UpsertDashboard boş Variables'ı "[]" yapıyor.
- Tablo ReplacingMergeTree — tam satır değişimi. Değişkenler gitti.
Belirti SESSİZ: hata yok, uyarı yok. VariablesBar çizilmiyor ve ${service} referansı taşıyan her panel satırı, substituteVars'ın "tanımsız değişkenli satırı düşür" kuralı gereği filtreden düşüyor — paneller sessizce FİLO GENELİNE geçiyor. Operatör bir servisi izlediğini sanırken tüm filoyu görüyor.
Bu, CLAUDE.md'nin ReplacingMergeTree değişmezinin doğrudan ihlali: "Whole-row replace: carry ALL fields forward." Kural yazılıydı, handler uymuyordu.
v0.9.559 — paylaşılan varlık-adı tarayıcısı.
Bir LLM yüzeyinin ürettiği SERBEST METİN, şema doğrulamasından geçse bile uydurma bir servis adı taşıyabilir: şema "string" der, içeriğe bakmaz. Enum'la kısıtlanan alanlar (root_cause.entity gibi) güvenlidir; ama özet, nedensellik zinciri ve önerilen aksiyon serbest metindir ve operatör asıl onları okur.
Somut kaçış: model `root_cause.entity` alanına beyaz listeden geçen gerçek bir servis yazar, sonra özette ve remediation.action'da HİÇ var olmayan bir servisi anlatır. Ekranda gerçek görünen bir nedensellik zinciri ve somut bir "şunu yeniden başlat" aksiyonu belirir. Enum kalkanı bunu görmez.
Tarayıcı YENİDEN YAZILMADI: copilot_aianalyze.go'daki postCheckServiceAnalysis bunu 2026'dan beri yapıyor. Buraya taşındı ki iki yüzey aynı kuralı paylaşsın — ikinci bir kopya yazmak, birinin diğerinden sessizce ayrışması demekti (bu oturumda tam bu sınıftan üç hata çıktı: v0.9.552/553/554).
v0.9.550 — evaluator sağlık okuması.
Operatör: "bazen sanki takıldığını hissediyorum." Bu his ölçülebilir değildi çünkü Problems sayfasının BOŞ olması iki bambaşka durumu birebir aynı gösteriyordu:
(a) her şey yolunda, açık problem yok (b) evaluator ölü/takılı, kimse problem üretmiyor
Bu endpoint o ikisini ayırır. Kalp atışını worker pod'u Redis'e yazar (evaluator/heartbeat.go), API pod'u buradan okur — dağıtık kurulumda evaluator ile /api/problems FARKLI pod'larda çalıştığı için paylaşımlı bir yerden geçmek zorunlu.
Tasarım kararı: TAZELİĞİ SUNUCU HESAPLAR. İstemciye ham zaman damgası verip "sen çıkar" demek, tarayıcı saati kayan bir makinede (VDI, uyku sonrası laptop) sessizce yanlış bir yaş üretirdi. Aynı dersi v0.9.543'te chNodeWork'te de uygulamıştık: geçen süre sunucunun generatedAt'inden türer.
v0.9.627 — patlama HIZI: exception önceliğinin eksik boyutu.
Operator-reported: tek servisten 12 dakikada 11.260 olay üreten bir exception grubu P2 göründü.
com.ibm.msg.client.jakarta.jms.DetailedInvalidDestinationException service bsa-cashmanagement-cashflow-prod first 04.08.2026 12:49:31 last 04.08.2026 13:01:36 toplam 11.260 → ~938 olay/dakika
exceptionPriority'nin P1 kapısı TEK bir koşuldu: "son 5 dakika içinde görülmüş VE toplam ≥ 500". Patlama 13:01'de bitti, operatör 13:22'de baktı — 21 dakikalık yaş kapıyı kapattı ve satır P2'ye düştü.
İki ayrı eksik vardı:
- TAZELİK PENCERESİ ÇOK DAR. P1 "şimdi ilgilen" demek; yirmi dakika önce biten 11 binlik bir patlama hâlâ şimdi ilgilenilmesi gereken bir şey. Olayın bitmesi etkisinin bittiği anlamına gelmiyor.
- HIZ HİÇ ÖLÇÜLMÜYORDU. 11.260 ile 110 arasındaki fark yalnız bir eşik karşılaştırmasına giriyordu, birim zamana değil. Üç günde birikmiş 11.260 ile on iki dakikada patlayan 11.260 aynı kovaya düşüyordu — halbuki ilki kronik, ikincisi olay.
Hız, elimizde ZATEN olan iki alandan türüyor: first_seen ve last_seen. v0.9.524'ün dersi burada da geçerli — sahip OLMADIĞIMIZ pencereli bir sayıyı uydurmuyoruz; "12 dakikada 11.260" birebir doğru bir cümle, "son 5 dakikada 11.260" ise yalan olurdu.
v0.9.553 — problem okuma zenginleştirmesi TEK çağrı noktası.
Operatör raporu: "Bazen P1 alertleri karıştırıyor ya da ben Problems sekmesinde chatte yazanları göremiyorum, orada da conflict var."
Kök sebep: "P1" bir ClickHouse kolonu değil, OKUMA ANI hesabı (chstore.computePriority). Hesabın kritik kolu RecentDeploy'a bakar:
postDeploy := p.RecentDeploy != nil && ... AgeSeconds <= 5*60 case postDeploy: return "P1", "critical + deploy Ns before" default: return "P2", "critical"
Yani deploy zenginleştirmesi koşmamışsa RecentDeploy nil kalır ve AYNI SATIR P2 olur. Sayfa yolları ikisini sırayla çağırıyordu, sohbet yolları (guided bundle'ları + vardiya özeti) yalnız priority'yi çağırıyordu — sonuç: taze deploy sonrası kritik bir problem sayfada P1, sohbette P2. Operatör iki yüzeyi yan yana koyduğunda çelişki görüyordu ve İKİSİ DE kendi içinde tutarlıydı, bu yüzden hangisinin yanlış olduğu belli olmuyordu.
Ayrışma yönü tek taraflıydı: sohbet ASLA sayfadan daha acil olamıyordu, hep daha sakin görünüyordu. Bu, en kötü yön — operatör sohbete güvenip acil bir şeyi ertelemiş olabilir.
Değişmez kural api.go'da zaten YAZILIYDI ("the deploys enrich must run before the priority enrich") ama yorum olarak; yorum derlenmez ve beş çağrı noktası onu ihlal etti. Bu dosya kuralı ÇAĞRILABİLİR hale getiriyor: iki adım tek fonksiyonda, sırası içeride sabit.
v0.9.559 — RCA kanıt kataloğu: İKİ ayrı kimlik uzayı.
Tasarım: docs/cosre-verdict-design.md §2.
Model, ürettiği her iddiayı bir kanıt kimliğine bağlamak zorunda. Ama kanıtın iki TÜRÜ var ve ikisi aynı işi göremez:
E1..En bulunmuş sinyaller → kök neden dayanağı OLABİLİR N1..Nn bakılmış, BULUNAMAMIŞ → yalnız ÇÜRÜTME dayanağı olabilir
Asimetri mantıksal, keyfi değil: "trafik artışı değil, çünkü istek hacmi sabit" geçerli bir çürütmedir; "pod restart döngüsünde, çünkü restart kaydı BULUNAMADI" saçmadır. Aranmış ve bulunamamış olmak, olduğunun kanıtı değildir.
Bu ayrım olmadan kalkan yalnız etkisiz kalmaz, ZARARLI olur: model negatif bir kaydı kanıt diye gösterir, kalkan geçer (kimlik gerçekten katalogdadır), sonra sunucu o katalog metnini iddianın yanına basar ve uydurma DOĞRULANMIŞ görünür. Kalkanın varlığı zararı artırır.
Katalog kurulurken ayrım VERİDEN yapılır (CheckedSignal.Found), modelin beyanından değil.
v0.9.559 — RCA impact: sayılar ClickHouse'tan, modelden DEĞİL.
Tasarım: docs/cosre-verdict-design.md §4-K5.
Modelden sayı istemek iki kere yanlış: uydurabilir, ve uydurduğunda operatör onu ÖLÇÜLMÜŞ sanır. Bu yüzden etki rakamları MV'den okunur ve modele yalnız YORUMLATILIR.
Ama "ClickHouse'tan geliyor" güvencesi, okuma yanlışsa daha tehlikelidir: operatör sayıyı sorgulamaz. Bu yüzden üç tuzak açıkça kapatıldı:
DOĞRU VARLIK. Birinci tasarım ankorun servisini okuyordu; tipik bir P1'de kök neden başka bir servistir ve "Kök neden: odeme-db" başlığının altında odeme-api'nin sayıları görünürdü.
BOŞ SONUÇ ≠ SIFIR. MV henüz materyalize olmamışsa (son ~5dk) sorgu 0 satır döner ve err == nil. Sıfır basmak "etkilenen istek yok" demek olurdu — ve ✨ Explain'e basılan an tam olarak olayın patladığı andır. Sayılar nil döner, sebep yazılır.
GELECEK PENCERE. boundAnalysisWindow ankorun başlangıcına 10dk ekler; taze bir problemde pencere sonu GELECEKTE kalır. Üst sınır now ile kırpılır, yoksa "sorgulanan aralık" yalan söyler.
Kova hizalaması ayrıca gerekiyordu ve v0.9.555'te mağaza katmanında düzeltildi (bu tasarımın çekişmeli incelemesinin yan ürünü).
v0.9.591 — verdict kaydının API tarafı.
Ayrı dosya çünkü sorumluluk ayrı: rca_verdict.go verdict'i ÜRETİR, bu dosya üretileni KAYDEDER. Kayıt yolu üretim yolunu asla etkilememeli ve bunu yapı gereği garanti etmenin en ucuz yolu iki dosyada tutmak.
v0.9.559 — RCA verdict deterministik kalkanları.
Tasarım: docs/cosre-verdict-design.md §4.
Temel ilke: model ne söylerse söylesin, SUNUCU doğrular. Şema biçimi zorlar, kalkanlar içeriği. Her kalkanın ne yakaladığı kadar NE YAKALAMADIĞI da yazılı — bir kalkanın sınırını bilmemek, olmamasından tehlikelidir.
v0.9.559 — RCA verdict orkestrasyonu.
Tasarım: docs/cosre-verdict-design.md
Akış: katalog kur → rakipleri üret → modele sor (şema kilitli) → KALKANLARDAN geçir → etkiyi ClickHouse'tan ölç → zarfla.
Modelin ürettiği ham şekil (rcaModelVerdict) ile operatöre dönen şekil (RCAVerdict) AYNI ŞEY DEĞİL: kalkanlar aradaki farkı üretir ve ne yaptıklarını rapor eder.
v0.9.665 — servis throughput'unu METRİKTEN okuma.
Operatör: "service overview throughput için ekrandaki metrikten job=abc/cm-put-service şeklinde Service ismi / sonrası olacak şekilde ayarlayabilir misin, metricten okusun bakalım bir. Oluştursun görelim."
Bugün Overview'ın throughput'u SPAN türevli (spanMetricBatch → service_summary_5m, giriş-span ilkesi). Bu uç ikinci bir kaynak veriyor: Prometheus biçimli bir sayaç metriği, servis kimliği `job` etiketinin son bölümünde.
TANILAMA UCUN ASIL İŞİ. Operatörün Grafana'sı PROMETHEUS'tan okuyor; o metriğin Coremetry'ye de girdiği garanti DEĞİL (collector onu OTLP ile iletiyor mu, adı korunuyor mu — buradan görülemez). Bu yüzden uç boş bir seri döndürüp susmuyor: metrik var mı, hangi `job` değerleri mevcut, desen ne — hepsini söylüyor. Böylece "veri yok" ile "desen tutmadı" ayırt edilebiliyor; boş bir grafik ikisini de aynı gösterirdi.
Package api handler for SLO autocreate (v0.5.147). Scans recent telemetry, picks the busiest N services, and stamps a baseline-grounded availability + latency SLO for each one the install doesn't already have. Admin-only, audit-logged, idempotent.
Tempo-compatible HTTP API. Lets Grafana add Coremetry as a Tempo datasource: only the endpoints Grafana actually calls during search and trace view are implemented (echo, search, search/tags, search/tag/{name}/values, traces/{id}).
Both the v1 and v2 paths are wired so older / newer Grafana versions work.
Spec reference: https://grafana.com/docs/tempo/latest/api_docs/
v0.9.632 — trace çözümlemesi TEK YERDE.
Operator-reported (prod, v0.9.631): Tempo fallback bir trace bulup waterfall'ı çizdiği hâlde "Explain trace" HTTP 404 "trace not found" veriyor.
Sebep bugünün tekrar eden sınıfı: kural iki yere bölünmüş ve ayrışmış. /api/traces/{id} handler'ı ClickHouse ıskalayınca Tempo'ya düşüyordu (api.go, "CH miss → Tempo fallback"); explain girdisini kuran buildTraceExplainInput ise DOĞRUDAN s.store.GetTrace çağırıyordu. Coremetry trace'i örneklemeyle dışarıda bıraktıysa CH'de sıfır span var — detay sayfası Tempo'dan okuyup 62 span gösteriyor, explain aynı trace için "yok" diyor.
Aynı boşluk BEŞ çağrı yerinde vardı: explain, trace paylaşım linki üretimi, paylaşılan snapshot görüntüleyici, span-detay ve trace karşılaştırma. Hepsi Tempo-only bir trace'te 404 veriyordu.
NEREYE UYGULANMAZ: internal/api/tempo.go'daki iki handler. Onlar Grafana'nın Tempo datasource'una Tempo API'si SUNUYOR — oradan yine Tempo'ya düşmek döngü olurdu.
traces_extras.go — the /api/traces?traceIds= phase-2-only path + the shared extraAttrs parsing (FAZ 2, docs/audit/traces-attribute-columns.md §6). Lives OUTSIDE api.go by operator constraint (api.go must not grow; the extraAttrs parse moved here from getTraces/exportTracesCSV for a net shrink).
No separate route is registered: `traceIds` is a param-branch of the existing GET /api/traces contract (the frontend's enrichment call rides the same client method family), and Go 1.22's ServeMux panics on a second "GET /api/traces" pattern — so getTraces delegates here first via serveTracesExtras and the api.go route table is untouched.
Contract:
GET /api/traces?traceIds=<id,id,…>&extraAttrs=<k,k…>&from=<ns>&to=<ns>
→ 200 {"extras": {"<traceId>": {"<key>": "<value>", …}, …}}
When traceIds is present, phase-1 (the list query) is SKIPPED entirely: only the bounded phase-2 (Store.TraceExtras — time-bounded WHERE + id IN-list + LIMIT) runs. from/to are REQUIRED — they are the phase-2 time bound; the client sends the visible rows' real min/max timestamps and the store pads `to` by the +5m slack. Without traceIds the request falls through to the normal getTraces flow, so the existing API contract (optional extraAttrs, extras inside each TraceRow) is unchanged.
Index ¶
- type AIRate
- type AnnotationItem
- type AnomalyRootCause
- type AvailablePage
- type CHAsyncInsert
- type CHClusterNode
- type CHCoordinator
- type CHCoordinatorSpread
- type CHHealth
- type CHMerge
- type CHMutation
- type CHNodeWork
- type CHNodeWorkResponse
- type CHPartHotspot
- type CHReplicationLag
- type CHSlowQuery
- type CHTopology
- type CacheKeyHit
- type CacheStatsSnapshot
- type CorrelationAnchor
- type CorrelationContext
- type CorrelationKind
- type CorrelationTrace
- type DeployRow
- type EvaluatorHealth
- type FleetRollout
- type FlowsResponse
- type GraphEdge
- type GraphNode
- type InboxAnomalyRef
- type InboxExceptionRef
- type InboxIncidentRef
- type InboxItem
- type InboxProblemRef
- type NoisyRuleWithSuggestion
- type RCAImpact
- type RCAVerdict
- type RootCause
- type SchemaColumn
- type SchemaTable
- type Server
- func (s *Server) EnableDemoMode(email, password string)
- func (s *Server) SetAutocomplete(a *acache.Store)
- func (s *Server) SetBackgroundConfig(b config.BackgroundConfig)
- func (s *Server) SetBuildVersion(v string)
- func (s *Server) SetCluster(c *cluster.Service)
- func (s *Server) SetLdapGroupSync(e *ldap.SyncEngine)
- func (s *Server) SetLockDegraded(b bool)
- func (s *Server) SetLogstoreESManager(m *logstore.ESManager)
- func (s *Server) SetMCP(m *mcp.Server)
- func (s *Server) SetPipeline(p *pipeline.Engine)
- func (s *Server) SetRAG(r *rag.Service)
- func (s *Server) SetRoles(ingest, apiRole bool)
- func (s *Server) SetTempo(t *tempo.Service)
- func (s *Server) SetThanos(t *thanos.Service)
- func (s *Server) SetVersion(v string)
- func (s *Server) Shutdown(ctx context.Context) error
- func (s *Server) Start() error
- func (s *Server) StartAuditDrainer(ctx context.Context)
- func (s *Server) StartCacheInvalidation(ctx context.Context)
- func (s *Server) StartRAGSync(ctx context.Context, lock cache.Lock)
- type ServiceGraphResponse
- type ServiceTopologyNode
- type ServiceTopologyResponse
- type SpanMetricServiceRow
- type TopologyNode
- type TopologyResponse
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type AIRate ¶ added in v0.5.167
type AIRate struct {
InputPer1M float64 `json:"inputPer1M"`
OutputPer1M float64 `json:"outputPer1M"`
}
AIRate is the per-model price quote used by the /ai cost estimate (v0.5.167). USD per 1M tokens, separately for input and output. Stored as a JSON blob in system_settings["ai.rates"] so admins can override the bundled defaults without code change. Local-model endpoints (Ollama, vLLM, LM Studio) should configure 0/0 — that's the default shape for any model not in the table.
type AnnotationItem ¶ added in v0.9.394
type AnnotationItem struct {
TS int64 `json:"ts"` // unix ns
Kind string `json:"kind"`
Title string `json:"title"`
Service string `json:"service,omitempty"`
TargetType string `json:"targetType,omitempty"` // problem | anomaly | event | rollout
TargetID string `json:"targetId,omitempty"`
Link string `json:"link,omitempty"`
}
AnnotationItem — şeridin tek olay modeli (beş marker mekanizmasının hedef birleşimi; Ş3 emeklilikleri bunun üstüne).
type AnomalyRootCause ¶ added in v0.8.167
type AnomalyRootCause struct {
RootCause
AnomalyID string `json:"anomalyId"`
AnomalyKind string `json:"anomalyKind"` // log_pattern | log_template_new | trace_op
Pattern string `json:"pattern"` // log pattern name OR operation name (trace_op)
}
AnomalyRootCause is the anomaly-anchored sibling of RootCause (v0.8.x). It embeds the SAME RootCause fan-out result so /anomalies and /problems share one rendering path later, and stamps the anchor from the AnomalyEvent (id/kind/pattern/service) instead of a Problem. Read-only.
type AvailablePage ¶ added in v0.5.251
type AvailablePage struct {
ID string `json:"id"` // route path; matches Sidebar href
Label string `json:"label"` // i18n key (e.g. "nav.inbox")
Group string `json:"group"` // i18n key for group heading
}
AvailablePage is one entry in the canonical sidebar-page registry that powers the custom-role checkbox grid in Settings → Roles. The frontend Sidebar.tsx mirrors these IDs as `href` values; keeping the registry on the backend means the role catalog's page IDs are validated against a single source of truth instead of drifting when somebody adds a sidebar entry.
IDs match the frontend route paths verbatim (e.g. "/inbox", "/services"). Group + label are i18n keys; the frontend resolves them via useT() so a language switch surfaces immediately.
AdminOnly pages are excluded from this registry — custom roles subset viewer's access only; admin/editor surfaces remain gated by their hard-coded role checks.
type CHAsyncInsert ¶ added in v0.5.346
type CHAsyncInsert struct {
Database string `json:"database"`
Table string `json:"table"`
TotalBytes uint64 `json:"totalBytes"`
EntriesCount uint64 `json:"entriesCount"`
FirstUpdateMsAgo uint64 `json:"firstUpdateMsAgo"`
}
CHAsyncInsert mirrors one row of system.asynchronous_inserts — a buffered INSERT awaiting the server-side coalescing flush. ageMs > 0 = how long it's been sitting; bytes/rows reflect what's queued so far. The whole list usually has a few rows during steady ingest, climbs visibly under burst (good sign — coalescence working) and falls back as the busy_timeout fires.
type CHClusterNode ¶ added in v0.5.388
type CHCoordinator ¶ added in v0.9.494
type CHCoordinator struct {
Host string `json:"host"`
// Initial — bu node'un GİRİŞ NOKTASI olduğu sorgu sayısı. Panelin
// asıl sinyali. query_log yolunda is_initial_query=1 sayımı,
// events yolunda ProfileEvent `InitialQuery`. İkisi aynı şeyi sayar.
Initial uint64 `json:"initial"`
Selects uint64 `json:"selects"`
Inserts uint64 `json:"inserts"`
Other uint64 `json:"other"`
ReadRows uint64 `json:"readRows"`
MemoryMB float64 `json:"memoryMB"`
// P50/P95 yalnız SELECT'ler üzerinden — INSERT süreleri karışırsa
// okuma gecikmesi okunamaz hale gelir. events yolunda 0 (o kaynakta
// süre dağılımı yok).
P50Ms float64 `json:"p50Ms"`
P95Ms float64 `json:"p95Ms"`
// UptimeS — yalnız events yolunda anlamlı. Sayaçlar sunucu
// açılışından beri biriktiği için, yeni restart etmiş bir node
// DÜŞÜK okunur; operatör dengesizliği uptime'a bakarak yorumlar.
UptimeS uint64 `json:"uptimeS,omitempty"`
}
Koordinatör dağılımı — hangi CH node'u KAÇ sorgunun giriş noktası oldu. "Giriş noktası" = system.query_log'da is_initial_query = 1 olan satır: Distributed bir sorguda yalnız koordinatör böyle loglar, shard'ların yaptığı alt-sorgular is_initial_query = 0'dır. Yani bu panel taramanın nerede yapıldığını DEĞİL, parse + fan-out + partial-state merge + final aggregation/sort/limit yükünün nerede biriktiğini ölçer.
Neden ayrı bir uç (v0.9.494): /admin/clickhouse gövdesi 5s TTL ile dönüyor ve içindeki query_log okuması yalnız > 500ms yavaş sorguları süzüyor. Buradaki okuma PENCEREDEKİ TÜM sorguları grupluyor — çok daha geniş bir satır kümesi. 5s'te bir koşturmak admin sayfasına gereksiz maliyet bindirir; teşhis panelinin 5 saniyelik tazeliğe ihtiyacı da yok. Kendi ucu + 60s TTL.
SELECT ve INSERT ayrı sayılıyor çünkü ikisi bugün FARKLI havuzlardan çıkıyor (internal/chstore/store.go:399-435):
- INSERT'ler ingest havuzunda, ConnOpenRoundRobin → 4 node'a yayılır (v0.9.481'in kazanımı; bu panelde dengeli görünmeli)
- SELECT'ler ana bağlantıda, ConnOpenInOrder → hep ilk node (açık kalem; okuma havuzu dilimi bunu düzeltecek)
Operatör tek tabloda "insert'lerim dağıldı, select'lerim dağılmadı" ayrımını görebilsin diye ikisi yan yana.
type CHCoordinatorSpread ¶ added in v0.9.494
type CHCoordinatorSpread struct {
Nodes []CHCoordinator `json:"nodes"`
Mode string `json:"mode"` // "cluster" | "standalone"
WindowS int64 `json:"windowS"` // ölçüm penceresi
// SelectImbalance / InsertImbalance = maxNode / ortalamaNode.
// 1.00 = kusursuz dağılım. N node'da tek node her şeyi alıyorsa
// değer N'e yaklaşır. Tek node'lu kurulumda 1.0 (anlamsız ama
// yanıltıcı değil) ve Mode = "standalone" ile birlikte okunur.
SelectImbalance float64 `json:"selectImbalance"`
InsertImbalance float64 `json:"insertImbalance"`
// InitialImbalance — panelin ASIL rakamı. Giriş-noktası sayısı en
// temiz koordinatör sinyali: shard'ların alt-sorguları buna girmez.
InitialImbalance float64 `json:"initialImbalance"`
// Source — "query_log" (pencereli, zengin) | "events" (açılıştan
// beri kümülatif, süre/satır dağılımı yok). Prod'da query_log
// çoğu kurulumda KAPALI (log_queries=0) ya da tablo hiç oluşmamış
// olur; panel o zaman sessizce boş kalmak yerine events'e düşer.
Source string `json:"source"`
// Note — panelin altında gösterilen insan-dili durum cümlesi.
// Boş değilse UI onu aynen basar; ölçüm yapılamadıysa NEDEN
// yapılamadığını söyler (boş panel + sessizlik yerine).
Note string `json:"note,omitempty"`
GeneratedAt int64 `json:"generatedAt"`
}
type CHHealth ¶ added in v0.5.319
type CHHealth struct {
// v0.5.388 — topology banner. First field on the payload so
// the operator's eye lands on the cluster mode + node count
// before drilling into the perf panels. Powered by:
// - cfg.ClusterName (operator-set env var) → "configured"
// side
// - system.clusters lookup (CH self-report) → "live" side
// - system.tables engine filter → which tables wear the
// Distributed wrapper vs plain MergeTree
// All three are needed because a misconfigured cluster name
// looks identical to standalone from the driver — only
// system.clusters confirms ZK actually wired the node up.
Topology CHTopology `json:"topology"`
SlowQueries []CHSlowQuery `json:"slowQueries"`
Merges []CHMerge `json:"merges"`
PartHotspots []CHPartHotspot `json:"partHotspots"`
Replication []CHReplicationLag `json:"replicationLag,omitempty"`
// v0.5.346 — in-flight async_insert batches. Each row =
// one currently-buffered INSERT awaiting flush. Lets the
// operator see whether the tuned async_insert_busy_timeout
// is doing useful coalescence or sitting idle.
AsyncInserts []CHAsyncInsert `json:"asyncInserts,omitempty"`
// v0.6.22 — in-flight ALTER TABLE … DELETE / UPDATE
// mutations. Healthy queue is empty; growing queue is an
// operator-visible signal that a state-table mutation
// pattern needs rethinking (tombstone, ReplacingMergeTree,
// etc.). Top 20 most-recent unfinished, sorted by parts-
// remaining desc so the worst offender lands on top.
Mutations []CHMutation `json:"mutations,omitempty"`
Generated int64 `json:"generatedAt"` // unix ns of snapshot
}
CHHealth — v0.5.319. Datadog-style ClickHouse dashboard payload. Gives the operator a single-page view of CH-side perf health: slow queries, merge queue depth, part overflow risk, async insert pressure. Powers the new /admin/clickhouse page.
Read directly from system.* tables which CH provides as built-in observability surface. Each panel is fail-isolated — a CH version without one of the views (e.g. older OS edition without system.async_inserts) leaves the slot at its zero value rather than 500-ing the whole page.
type CHMerge ¶ added in v0.5.319
type CHMerge struct {
// Host (v0.9.540) — merge'i KOŞTURAN node. Küme yapılandırılmamışsa
// bağlı olunan node'un adı. 4 node'lu kümede "hangi node merge'e
// boğulmuş" sorusunun cevabı bu kolonda.
Host string `json:"host"`
Database string `json:"database"`
Table string `json:"table"`
ElapsedSec float64 `json:"elapsedSec"`
ProgressPct float64 `json:"progressPct"`
RowsRead uint64 `json:"rowsRead"`
MergedSize uint64 `json:"mergedSizeBytes"`
}
type CHMutation ¶ added in v0.6.22
type CHMutation struct {
Database string `json:"database"`
Table string `json:"table"`
Command string `json:"command"` // trimmed; full text would be unbounded
Parts uint64 `json:"parts"` // parts left to mutate
ElapsedMs int64 `json:"elapsedMs"` // since the mutation was submitted
LatestFail string `json:"latestFail,omitempty"`
}
CHMutation — v0.6.22. One row per in-flight or recently-stuck ALTER TABLE … DELETE / UPDATE mutation. ALTER mutations rewrite matching parts in the background; healthy systems clear them within seconds. A growing queue is the operator-visible signature that a state table is being mutated faster than CH can rebuild parts — at which point the right fix is a tombstone or ReplacingMergeTree pattern, not "wait longer".
Powered by system.mutations. Filter is_done = 0 so only the queue depth shows up; finished mutations age out of the table after ~7 days and don't matter for live ops.
type CHNodeWork ¶ added in v0.9.543
type CHNodeWork struct {
Host string `json:"host"`
// Shard/Replica — system.clusters'tan, node'un KENDİ kaydı
// (is_local). Dengesizliği doğru okumak için şart: aynı shard'ın
// replikaları AYNI veriyi tutar, farklı shard'lar tanım gereği
// farklı. Shard etiketi olmadan shard çarpıklığı "merge suçu" gibi
// okunur. 0 = küme tanımlı değil ya da eşleşme bulunamadı.
Shard uint32 `json:"shard"`
// Replica — makro bir AD döndürür ("chc-0"), sayı değil. Gösterim
// için string; gruplama zaten SHARD üzerinden yapılıyor.
Replica string `json:"replica"`
// UptimeS — sayaçların kapsadığı süre. İstemci farkı bununla
// doğrular: uptime GERİYE giderse node yeniden başlamıştır ve
// taban geçersizdir (fark 0'a kelepçelenirse node "iş yapmıyor"
// görünür — mümkün en yanıltıcı değer).
UptimeS uint64 `json:"uptimeS"`
// Ham kümülatif sayaçlar. İstemci fark alır.
CPUMicros uint64 `json:"cpuMicros"` // OSCPUVirtualTimeMicroseconds
MergeMillis uint64 `json:"mergeMillis"` // MergeExecuteMilliseconds
InsertedRows uint64 `json:"insertedRows"` // InsertedRows
PartFetches uint64 `json:"partFetches"` // ReplicatedPartFetches
SelectedBytes uint64 `json:"selectedBytes"` // SelectedBytes
MergesLaunched uint64 `json:"mergesLaunched"` // Merge (events → kümülatif)
}
type CHNodeWorkResponse ¶ added in v0.9.543
type CHNodeWorkResponse struct {
Nodes []CHNodeWork `json:"nodes"`
Mode string `json:"mode"` // "cluster" | "standalone"
// Source — "events" (okundu) | "none" (okunamadı). Ayrım şart:
// boş tablo "her şey dengeli" diye okunmamalı.
Source string `json:"source"`
// ShardsKnown — shard etiketleri çözülebildi mi. False ise UI
// dengesizliği shard-içi hesaplayamaz ve bunu SÖYLEMELİ.
ShardsKnown bool `json:"shardsKnown"`
// ExpectedNodes — system.clusters'ın bildirdiği node sayısı.
// len(Nodes) bundan azsa bir node yanıt vermemiştir; eksik küme
// "dengeli" okunursa tam da aradığımız node gözden kaçar.
ExpectedNodes int `json:"expectedNodes"`
Note string `json:"note,omitempty"`
GeneratedAt int64 `json:"generatedAt"`
}
type CHPartHotspot ¶ added in v0.5.319
type CHReplicationLag ¶ added in v0.5.319
type CHSlowQuery ¶ added in v0.5.319
type CHTopology ¶ added in v0.5.388
type CHTopology struct {
// Mode is "cluster" when an ON CLUSTER name is configured AND
// system.clusters confirms it exists; "standalone" otherwise.
// A misconfigured cluster name (operator set the env var but
// the CH server has no matching <remote_servers> block) shows
// up as "standalone" with a non-empty ConfiguredCluster — the
// banner flags this mismatch.
Mode string `json:"mode"`
ConfiguredCluster string `json:"configuredCluster,omitempty"`
Database string `json:"database"`
ConnectedHosts []string `json:"connectedHosts,omitempty"`
// Nodes — one entry per (shard, replica) registered in
// system.clusters for the configured cluster. Empty in
// standalone mode.
Nodes []CHClusterNode `json:"nodes,omitempty"`
// Table engine breakdown — drives the "distributed table?"
// answer. DistributedTables > 0 = the install is running the
// full cluster pattern (Distributed wrapper + _local
// Replicated tables). LocalReplicated = Replicated*MergeTree
// count. Plain = MergeTree / ReplacingMergeTree without
// replication. Used to confirm migrations actually built the
// cluster pattern they were supposed to.
DistributedTables int `json:"distributedTables"`
LocalReplicated int `json:"localReplicated"`
PlainMergeTree int `json:"plainMergeTree"`
// ZK / Keeper presence — system.zookeeper exists only when
// CH is configured with a Keeper endpoint. ReplicatedMergeTree
// needs this; absence on a cluster mode install is a config
// bug that should surface here, not via a CREATE TABLE failure
// six hours into ingest.
ZooKeeperConnected bool `json:"zookeeperConnected"`
// v0.5.419 — resolved per-table shard expression map. Lets
// the operator audit "which shard key did each table actually
// get?" without `SHOW CREATE TABLE` round-trips. Empty in
// standalone mode.
ShardPolicy map[string]string `json:"shardPolicy,omitempty"`
// v0.5.428 — bug-fix: distinguish "system.clusters probe failed"
// from "probe succeeded but returned no rows". Without this the
// frontend's misconfig banner false-positives when CH is busy
// enough to time out the probe (v0.5.424 raised the timeout
// from 3s to 8s, but operators still hit it intermittently).
// Empty when the probe completed cleanly (regardless of result
// count); populated with the error string when it failed. The
// frontend treats a populated error as "soft warning, can't
// confirm cluster" rather than the hard "misconfig" banner.
ClusterProbeError string `json:"clusterProbeError,omitempty"`
// v0.5.439 — when the live probe fails but a recent successful
// probe is cached, Nodes is filled from the cache and these
// two flags surface the staleness so the frontend renders a
// small "last refreshed N min ago" pill instead of the warn
// banner. ClusterProbeError stays empty in this path — the
// banner only fires when there's no cache to fall back on.
ClusterNodesStale bool `json:"clusterNodesStale,omitempty"`
ClusterNodesAgeMs int64 `json:"clusterNodesAgeMs,omitempty"`
}
CHTopology — what cluster does Coremetry think it's talking to, and does the live CH agree.
type CacheKeyHit ¶ added in v0.5.37
type CacheStatsSnapshot ¶ added in v0.5.37
type CacheStatsSnapshot struct {
SinceUnixNano int64 `json:"sinceUnixNano"`
Counts map[string]int64 `json:"counts"`
TopKeys []CacheKeyHit `json:"topKeys"`
L1Size int `json:"l1Size"`
L1Cap int `json:"l1Cap"`
}
CacheStatsSnapshot is the wire shape for the admin endpoint. Counts is a tier → hit count map; TopKeys is a sorted slice of the most-frequently-served keys (capped at 20).
type CorrelationAnchor ¶ added in v0.8.153
type CorrelationAnchor struct {
Kind CorrelationKind `json:"kind"`
TraceID string `json:"traceId,omitempty"`
Service string `json:"service,omitempty"`
TsNs int64 `json:"tsNs,omitempty"`
FromNs int64 `json:"fromNs"`
ToNs int64 `json:"toNs"`
// JoinKey is the strongest join the bundle actually used:
// "trace_id" — exact cross-signal join (no time fuzz)
// "exemplar" — a real representative trace from the spanmetrics
// rollup (or the slowest/erroring raw span): the
// metric→trace pivot for latency/error anchors. Not the
// literal data point, but exact-enough to pivot into.
// "service+window" — the genuinely-fuzzy fallback (throughput/count
// metric, or no exemplar resolved): the operator must
// see this.
JoinKey string `json:"joinKey"`
}
CorrelationAnchor is what the operator pivoted FROM, echoed back so the drawer's anchor header + join-key chip can render which join is being trusted.
type CorrelationContext ¶ added in v0.8.153
type CorrelationContext struct {
Anchor CorrelationAnchor `json:"anchor"`
Trace *CorrelationTrace `json:"trace,omitempty"`
Logs []*logstore.LogRecord `json:"logs"` // always present (possibly empty)
Metrics []chstore.SpanMetricSeries `json:"metrics"` // anchor service RED series (possibly empty)
Exemplar *chstore.Exemplar `json:"exemplar,omitempty"` // metric anchor: a REAL representative trace to pivot INTO (slow/error rollup exemplar, raw-span fallback)
}
CorrelationContext is the assembled pivot bundle. Every lens is best-effort: a lens with no data soft-fails to nil/empty, exactly like rootcause.go — a partial bundle still helps the operator.
type CorrelationKind ¶ added in v0.8.153
type CorrelationKind string
CorrelationKind is the pivot anchor's signal shape.
const ( CorrelateTrace CorrelationKind = "trace" CorrelateLog CorrelationKind = "log" CorrelateMetric CorrelationKind = "metric" )
type CorrelationTrace ¶ added in v0.8.153
type CorrelationTrace struct {
TraceID string `json:"traceId"`
RootName string `json:"rootName"`
Service string `json:"service"`
DurationMs float64 `json:"durationMs"`
SpanCount int `json:"spanCount"`
Services []string `json:"services"`
ErrSpans int `json:"errSpans"`
StartTimeNs int64 `json:"startTimeNs"`
EndTimeNs int64 `json:"endTimeNs"`
// Spans is the raw span list (capped) so the drawer's extracted
// ServiceTimeline sub-component renders the SAME per-service density bars
// TracePeekDrawer does — no second derivation. Capped to keep the bundle
// bounded for very large traces.
Spans []chstore.SpanRow `json:"spans"`
}
CorrelationTrace is the condensed trace lens — enough for the service-timeline mini-waterfall + a header, without re-loading the full /trace page. Derived from the same GetTrace spans the /api/traces/{id} endpoint returns.
type DeployRow ¶ added in v0.9.436
type DeployRow struct {
chstore.RecentDeployEntry
Impact *chstore.DeployImpact `json:"impact,omitempty"`
}
DeployRow (v0.9.436) — deploy girdisi + opsiyonel etki deltası.
type EvaluatorHealth ¶ added in v0.9.550
type EvaluatorHealth struct {
// Status — dört hal, dördü de FARKLI şeyler söyler:
//
// ok — son tik taze ve temiz
// stale — kalp atışı var ama bayat (takılma / lider yok /
// worker dağıtılmamış)
// failing — son tik hata ile bitti (deadline dahil)
// unknown — ÖLÇEMİYORUZ (Redis yok veya hiç kalp atışı
// yazılmamış). "ok" DEĞİLDİR ve öyle gösterilmemeli;
// bilmemeyi iyi habere çevirmek bu işin tam da
// düzeltmeye çalıştığı hata.
Status string `json:"status"`
// Reason — operatöre gösterilecek tek cümlelik açıklama.
Reason string `json:"reason"`
// AgeSec — son tikin bitişinden bu yana geçen süre (sunucu saati).
// Status=unknown iken anlamsızdır, -1 döner.
AgeSec int64 `json:"ageSec"`
// Son tikin ayrıntıları; unknown iken sıfır değerli.
DurationMS int64 `json:"durationMs"`
Rules int `json:"rules"`
Opened int `json:"opened"`
Resolved int `json:"resolved"`
Err string `json:"err,omitempty"`
Version string `json:"version,omitempty"`
}
EvaluatorHealth — /api/problems/evaluator yanıtı.
type FleetRollout ¶ added in v0.9.435
FleetRollout — Rollout + hangi servise ait olduğu (filo satırı).
type FlowsResponse ¶ added in v0.5.103
type FlowsResponse struct {
Flows []chstore.RootFlow `json:"flows"`
From int64 `json:"from"`
To int64 `json:"to"`
// v0.7.39 — total distinct flows in the window (the list is capped at
// ?top). >len(Flows) → the UI shows "showing N of M flows — raise top".
TotalFlows int `json:"totalFlows,omitempty"`
}
FlowsResponse lists the top root-anchored business flows in a window. Each entry pairs a root signature with its trace count and the unique set of services those traces touched.
type GraphEdge ¶ added in v0.8.10
type GraphEdge struct {
Source string `json:"source"`
Target string `json:"target"`
Calls uint64 `json:"calls"`
Errors uint64 `json:"errors"`
ErrorRate float64 `json:"errorRate"`
Rate float64 `json:"rate"` // calls per minute over the window (v0.8.x)
AvgMs float64 `json:"avgMs"`
P99Ms float64 `json:"p99Ms"`
Protocol string `json:"protocol,omitempty"` // http | grpc | db | kafka — SpanKind proxy
}
GraphEdge is one directed caller→callee edge carrying RED metrics + protocol.
type GraphNode ¶ added in v0.8.10
type GraphNode struct {
ID string `json:"id"` // canonical id (the MV's raw name, e.g. "payments" or "db:h2")
Name string `json:"name"` // display name, prefix-decoded ("payments", "h2")
Kind string `json:"kind"` // service | database | queue | external | internal
System string `json:"system,omitempty"` // db.system / messaging.system when applicable
DbName string `json:"dbName,omitempty"` // db.name (schema/instance) — database nodes only
Env string `json:"env,omitempty"` // deployment.environment
Calls uint64 `json:"calls"` // node throughput (inbound preferred, else outbound)
Errors uint64 `json:"errors"`
ErrorRate float64 `json:"errorRate"` // (errors/calls)*100 — drives health color
Rate float64 `json:"rate"` // calls per minute over the window — node-size encoding (v0.8.x)
// v0.9.367 — which population Calls/Errors/ErrorRate came from:
// "inbound" (calls INTO the node — the normal basis) or "outbound"
// (the entry-service fallback: no instrumented caller, so the totals
// are what the node's DEPENDENCIES returned). The UI labels the two
// differently; without this the fallback was silent and a gateway's
// health dot read as its own error rate.
CallsBasis string `json:"callsBasis,omitempty"`
}
GraphNode is one node in the OTel-native service map.
type InboxAnomalyRef ¶ added in v0.5.211
type InboxExceptionRef ¶ added in v0.5.211
type InboxIncidentRef ¶ added in v0.9.321
type InboxIncidentRef struct {
ID string `json:"id"`
Severity string `json:"severity"`
Status string `json:"status"`
}
InboxIncidentRef — v0.9.321. A declared Incident is the one triage object a HUMAN created on purpose, and it was the only source the merged queue never showed: an operator working from /inbox could miss an open incident entirely while the sidebar's own /incidents badge counted it.
type InboxItem ¶ added in v0.5.211
type InboxItem struct {
ID string `json:"id"` // composite: "<kind>:<nativeId>"
Kind string `json:"kind"` // problem | exception | anomaly
Source string `json:"source"` // human label: "Alert rule" / "Exception" / "Anomaly"
Priority string `json:"priority"` // P1 | P2 | P3
PriorityReason string `json:"priorityReason"`
Severity string `json:"severity"` // critical | warning | info
Service string `json:"service"`
Title string `json:"title"` // rule name / exception type / pattern
Description string `json:"description"`
StartedAt int64 `json:"startedAt"` // unix ns
LastSeen int64 `json:"lastSeen"` // unix ns; for problems == StartedAt
Assignee string `json:"assignee,omitempty"`
// OwnerTeam + SRETeam attached server-side from
// service_metadata so the inbox can render team chips
// without each row firing a per-service lookup. Empty when
// no catalog row exists for the service. OwnerTeam mirrors
// what's auto-set on Problem.Assignee at open time;
// surfacing it on every row (even exceptions / anomalies)
// keeps the column meaningful across kinds.
OwnerTeam string `json:"ownerTeam,omitempty"`
SRETeam string `json:"sreTeam,omitempty"`
Status string `json:"status"` // open | acknowledged | resolved (problems);
// open | regressed (exceptions); active | cleared (anomalies)
Clusters []string `json:"clusters,omitempty"`
// v0.9.255 — enrichment results the inbox was already PAYING for and
// then dropping. listInbox runs EnrichProblemsWithRunbooks /
// WithDeploys before mapping (three CH round-trips per poll), but
// problemToInbox never copied the results out, so every triage row
// arrived without the two facts an operator reaches for first:
// "is there a runbook" and "did something just deploy". The queries
// were billed and the answers thrown away.
//
// Problem-kind only for now: exceptions and anomalies have their own
// deploy correlation paths and are not enriched on this route.
RunbookURL string `json:"runbookUrl,omitempty"`
RecentDeploy *chstore.RecentDeploy `json:"recentDeploy,omitempty"`
// AISummary (v0.9.530) — AYNI hata sınıfının ikinci nüshası. Hem
// Problem.AISummary (problem.go:789 SELECT'inde) hem
// ExceptionGroup.AISummary (exception_inbox.go SELECT'inde) bu
// handler'ın belleğine ZATEN geliyordu; mapper ikisini de atıyordu.
// Faturası ödenmiş, cevabı çöpe atılmış — yukarıdaki v0.9.255
// yorumunun tarif ettiği durumun aynısı, 1300 satır aşağıda.
//
// Sunucuda kırpılır: özet tek cümle değil, çok bölümlü bir blok
// ("Olası neden: / Kanıt: / İlk kontroller:"), ~700 karakter. Tam
// metnin yeri detay yüzeyi; satırın işi TARAMA.
//
// AISummaryAt olmadan gönderilmez: özet tek yazımlıktır ama satırın
// gövdesi (occurrences, mesaj) altından değişmeye devam eder, ve
// yaşsız bir çıkarım canlı sayının altında taze görünür.
AISummary string `json:"aiSummary,omitempty"`
AISummaryAt int64 `json:"aiSummaryAt,omitempty"`
// Kind-specific drill-down hints. Only one is populated per
// row. Keeps the JSON shape skinny — frontend reads exactly
// the one matching `kind`.
Problem *InboxProblemRef `json:"problem,omitempty"`
Exception *InboxExceptionRef `json:"exception,omitempty"`
Anomaly *InboxAnomalyRef `json:"anomaly,omitempty"`
Incident *InboxIncidentRef `json:"incident,omitempty"`
}
InboxItem is the unified shape every triage-worthy thing (Problem / Exception group / Anomaly event) collapses into. Kind discriminates which source it came from; the kind- specific blob carries the bits needed to drill-down.
Designed so a single table on /inbox can show "everything needing a human" without operators tab-hopping between Problems / Exceptions / Anomalies pages — same priority blend, same age column, same assignee column. The per-source pages still exist as drill-down targets.
type InboxProblemRef ¶ added in v0.5.211
type NoisyRuleWithSuggestion ¶ added in v0.5.131
type NoisyRuleWithSuggestion struct {
chstore.NoisyRule
Suggestion string `json:"suggestion"`
SuggestedFor uint32 `json:"suggestedForSec,omitempty"`
SuggestedMin uint32 `json:"suggestedMinSamples,omitempty"`
SuggestedCD uint32 `json:"suggestedCooldownSec,omitempty"`
CurrentFor uint32 `json:"currentForSec"`
CurrentMin uint32 `json:"currentMinSamples"`
CurrentCD uint32 `json:"currentCooldownSec"`
}
NoisyRuleWithSuggestion enriches the raw NoisyRule with a heuristic suggestion string + structured deltas the UI can apply with one click via the existing AlertRule edit endpoint.
type RCAImpact ¶ added in v0.9.559
type RCAImpact struct {
// Entity — sayıların AİT OLDUĞU varlık. Ankor servisi olmayabilir;
// bu yüzden alan açıkça taşınıyor.
Entity string `json:"entity"`
// AnchorService — problemin/anomalinin kendi servisi. Ayrı tutulur
// ki operatör hangi sayının neye ait olduğunu karıştırmasın.
AnchorService string `json:"anchorService,omitempty"`
// nil = ÖLÇEMEDİK (sıfır DEĞİL).
ErrorCount *uint64 `json:"errorCount"`
RequestCount *uint64 `json:"requestCount"`
WindowFromNs int64 `json:"windowFromNs"`
WindowToNs int64 `json:"windowToNs"`
// Note — ölçüm ölçülemediyse SEBEBİ. Boş = sayılar geçerli.
Note string `json:"note,omitempty"`
}
RCAImpact — ölçülmüş etki. Ölçülemeyen alanlar nil.
type RCAVerdict ¶ added in v0.9.559
type RCAVerdict struct {
Verdict string `json:"verdict"`
Title string `json:"title"`
Summary string `json:"summary"`
RootCause rcaModelRootCause `json:"rootCause"`
CausalChain []rcaModelChainStep `json:"causalChain,omitempty"`
RejectedHypotheses []rcaModelRejected `json:"rejectedHypotheses,omitempty"`
MissingEvidence []string `json:"missingEvidence,omitempty"`
Remediation []rcaModelRemediation `json:"remediation,omitempty"`
// Confidence — TAVANLANMIŞ nihai güven (bkz. capRCAConfidence).
Confidence float64 `json:"confidence"`
// ModelConfidence — modelin kendi BEYANI. Ayrı tutulur ki tavanın
// ne kadar indirdiği görünsün.
ModelConfidence float64 `json:"modelConfidence"`
// HypothesisConfidence — deterministik korelasyon motorunun güveni.
//
// Üç ayrı "confidence" aynı ekranda buluşuyordu ve üçü farklı
// şeydi; adlandırma bu yüzden açık (tasarım §3).
HypothesisConfidence float64 `json:"hypothesisConfidence"`
// Evidence — modelin atıf yaptığı kanıtların SUNUCU metni.
// Modelin ürettiği hiçbir metin buraya girmez.
Evidence []rcaEvidenceRef `json:"evidence,omitempty"`
Impact *RCAImpact `json:"impact,omitempty"`
Shields rcaShieldReport `json:"shields"`
}
RCAVerdict — kalkanlardan geçmiş verdict.
type RootCause ¶ added in v0.7.51
type RootCause struct {
ProblemID string `json:"problemId"`
Service string `json:"service"`
Metric string `json:"metric"`
StartedAt int64 `json:"startedAt"`
FromNs int64 `json:"fromNs"`
ToNs int64 `json:"toNs"`
RecentDeploy *chstore.RecentDeploy `json:"recentDeploy,omitempty"`
Correlations []chstore.ChangedService `json:"correlations"`
BlastRadius *chstore.BlastRadius `json:"blastRadius,omitempty"`
BubbleUp *chstore.BubbleUpResult `json:"bubbleUp,omitempty"`
Exemplar *chstore.Exemplar `json:"exemplar,omitempty"`
}
RootCause is the assembled "what changed / likely cause" bundle for one Problem (v0.7.51). It orchestrates signals that already exist but were scattered — recent deploy, correlated service changes, dimension bubble-up, blast radius, an exemplar trace — into a single cached read so the Problem triage drawer shows one root-cause surface instead of the operator hopping across pages. Read-only.
type SchemaColumn ¶
type SchemaTable ¶
type SchemaTable struct {
Table string `json:"table"`
Engine string `json:"engine"`
Columns []SchemaColumn `json:"columns"`
}
SchemaTable + SchemaColumn drive the playground's left-side browser. Operator clicks a column → it pastes into the editor at cursor; click a table → SELECT * FROM <table> LIMIT 100 gets generated. Both query system.tables / system.columns, pre-filtered to the current database. Results cached 60s.
type Server ¶
type Server struct {
// contains filtered or unexported fields
}
func (*Server) EnableDemoMode ¶
EnableDemoMode wires the demo credentials returned by /api/auth/config. Loud no-op when called with empty credentials so a misconfigured demo flag doesn't silently expose nothing.
func (*Server) SetAutocomplete ¶ added in v0.8.80
SetAutocomplete wires the Redis autocomplete cache (v0.8.80). Called once from main() after the api.Server is constructed. nil-safe — the picker handlers fall back to ClickHouse when it's absent or cold.
func (*Server) SetBackgroundConfig ¶ added in v0.4.95
func (s *Server) SetBackgroundConfig(b config.BackgroundConfig)
SetBackgroundConfig wires the cadence/timeout knobs to the Server. Called once from main() after Load() so the status probe respects the configured ceiling.
func (*Server) SetBuildVersion ¶ added in v0.9.339
SetBuildVersion records what the IMAGE actually is, independent of any display override. v0.5.394 removed the env override precisely because a stale value masked this; keeping both is what lets the override come back safely (v0.9.339).
func (*Server) SetCluster ¶ added in v0.5.253
SetCluster wires the per-pod heartbeat / membership service (v0.5.253). Always called from main() — the service degenerates to a single-pod view when Redis isn't configured so handlers don't need to nil-check before calling Members.
func (*Server) SetLdapGroupSync ¶ added in v0.8.526
func (s *Server) SetLdapGroupSync(e *ldap.SyncEngine)
SetLdapGroupSync wires the LDAP group-sync engine (v0.8.526).
func (*Server) SetLockDegraded ¶ added in v0.8.212
SetLockDegraded records that the distributed leader lock fell back to the always-leader Noop despite Redis being configured (Redis down at boot) — so /admin/stats can warn that multi-pod background jobs are duplicated. v0.8.212. Called with false by the Redis re-probe (v0.8.341) once the real lock is swapped back in — the /admin/stats warning clears without a pod restart.
func (*Server) SetLogstoreESManager ¶ added in v0.8.232
SetLogstoreESManager wires the UI-managed logstore config owner (v0.8.232). main() constructs the manager alongside the Switchable logstore; the Settings → Elasticsearch handlers are 503 no-ops without it (partial init / tests).
func (*Server) SetMCP ¶ added in v0.6.4
SetMCP wires the Model Context Protocol server (v0.6.4). Called once from main() after the api.Server is constructed. nil is valid — leaves the /api/mcp/* routes unregistered.
func (*Server) SetPipeline ¶ added in v0.5.263
SetPipeline wires the engine. Always called from main(); the admin handlers nil-check so a misconfigured boot doesn't 500 every request — they return 503 with a clear reason.
func (*Server) SetRAG ¶ added in v0.8.441
SetRAG bağlar (v0.8.438) — cluster.Set deseninde opsiyonel bağımlılık.
func (*Server) SetRoles ¶ added in v0.8.346
SetRoles wires the pod's runtime role split into the HTTP surface (v0.8.346, HA audit H6). main.go's old comment claimed "api.NewServer handles the role guard internally" — no such code existed: every role registered POST /v1/* while only ingest pods Start() the consumers, so a collector pointed at an api-role pod had its Exports 200-OK'd into channels NOBODY DRAINED (silent black hole; queue gauges even looked healthy at a constant 100%). Defaults (unset) = all roles on, which keeps monolithic mode and test-constructed Servers byte-identical.
func (*Server) SetTempo ¶ added in v0.5.208
SetTempo wires the external Tempo client. Always called from main() with a non-nil service — Configured() reports whether the operator has actually filled in the settings.
func (*Server) SetThanos ¶ added in v0.8.576
SetThanos wires the multi-cluster Thanos Querier client (v0.8.576). Always called from main() with a non-nil service — HasEnabledClusters() reports whether the operator configured any cluster.
func (*Server) SetVersion ¶
SetVersion records the build-time tag. Called once from main(); safe to call before Start() since /api/version is only consulted by SPA after the server is listening.
func (*Server) Shutdown ¶ added in v0.8.336
Shutdown drains the HTTP server (v0.8.336, HA audit H1): stops accepting new connections, lets in-flight requests finish within ctx's deadline. Safe before Start() (nil server = no-op).
func (*Server) StartAuditDrainer ¶ added in v0.5.339
StartAuditDrainer runs the batched audit-write loop until ctx is cancelled. Triggers a flush when either the channel hits 64 pending entries or the 200ms tick elapses — whichever comes first. Errors are logged but don't tear down the drainer; the next tick reattempts.
func (*Server) StartCacheInvalidation ¶ added in v0.5.337
StartCacheInvalidation subscribes to the invalidation channel and drains incoming messages into l1.del. Runs once per Server; the subscription lifetime is bound to the server's lifetime context. When Subscribe returns an error (Redis down, or pub/sub unsupported), we log and exit — the soft TTL keeps the L1 tier from growing stale unbounded, just for longer.
Called from main.go alongside the other StartConfigRefresh loops, exported because the constructor doesn't take a ctx.
type ServiceGraphResponse ¶ added in v0.8.10
type ServiceGraphResponse struct {
Nodes []GraphNode `json:"nodes"`
Edges []GraphEdge `json:"edges"`
Scope string `json:"scope"`
Focus string `json:"focus,omitempty"`
// TotalNodes / ShownNodes (v0.8.295, re-land of v0.8.277) — set by
// pruneServiceGraphTopN. When the global render budget trims a large
// graph, ShownNodes < TotalNodes and the UI can show "showing X of Y
// services" (same contract as the v0.8.215 cap on sampled
// /api/service-map).
TotalNodes int `json:"totalNodes"`
ShownNodes int `json:"shownNodes"`
}
ServiceGraphResponse is the compact payload the canonical renderer consumes.
type ServiceTopologyNode ¶ added in v0.5.102
type ServiceTopologyNode struct {
ID string `json:"id"` // canonical id used by edges (service name OR "db:postgresql")
Name string `json:"name"` // display label, sans prefix for infra ("postgresql" not "db:postgresql")
Kind string `json:"kind"` // "service" | "db" | "queue" | "external"
// v0.5.312 — Phase 2 enrichment for the topology redux:
// soft-cluster the diagram by k8s.namespace / service.namespace
// and paint each node with a health badge from open-problems
// count. Both are read-time-enriched (no schema change), nil-
// safe (omitempty), so older frontends keep working.
Namespace string `json:"namespace,omitempty"`
Health string `json:"health,omitempty"` // "" | "green" | "yellow" | "red"
HealthReason string `json:"healthReason,omitempty"` // short "2 open criticals" etc.
OpenCritical int `json:"openCritical,omitempty"`
OpenWarning int `json:"openWarning,omitempty"`
// v0.5.409 — known 3rd-party SaaS / cloud annotation for
// external nodes. Populated from the edge's ExtDisplay /
// ExtKind (set by external_catalogue lookup). UI renders a
// human-readable display name + category badge instead of
// the raw `ext:api.stripe.com` hostname.
ExtDisplay string `json:"extDisplay,omitempty"`
ExtKind string `json:"extKind,omitempty"`
// v0.5.410 — display-only environment annotation
// (deployment.environment / service.namespace /
// k8s.namespace.name). Populated from the edge that
// brought this node into the graph. UI surfaces it as a
// small chip ("prod" / "stage") next to the service name
// so multi-env installs distinguish at-a-glance.
Env string `json:"env,omitempty"`
// v0.7.32 — for a collapsed broadcast queue node, the real number of
// distinct consumer services its fan-out was hidden behind. The renderer
// shows "→ N services (broadcast)" on the node instead of N edges. Only set
// on queue nodes whose consumer count exceeded the broadcast threshold.
BroadcastFanout int `json:"broadcastFanout,omitempty"`
}
ServiceTopologyNode is one node in the service-level graph. Kind distinguishes a real service from synthetic infra nodes (db, queue, cache, external) so the renderer can paint them differently without per-node lookup.
type ServiceTopologyResponse ¶ added in v0.5.102
type ServiceTopologyResponse struct {
Nodes []ServiceTopologyNode `json:"nodes"`
Edges []chstore.ServiceTopologyEdge `json:"edges"`
From int64 `json:"from"`
To int64 `json:"to"`
Truncated bool `json:"truncated"`
// v0.6.48 — server-side scoping for thousand-service fabrics.
// TotalServices is the distinct service count BEFORE the top-N /
// focus bound was applied, so the UI can show "showing N of M
// services — search or focus to refine". Scoped is true when the
// returned graph is a bounded subset (top-N by call volume, or a
// focus neighbourhood) rather than the full fabric.
TotalServices int `json:"totalServices"`
Scoped bool `json:"scoped"`
ScopeReason string `json:"scopeReason,omitempty"` // "top-50 by call volume" | "focus: <svc> +2 hops"
// v0.7.32 — number of broadcast queue topics whose consumer fan-out was
// collapsed (a topic with >threshold distinct consumers, e.g. a kafka
// cache-refresh broadcast). The UI shows a "N broadcast topics collapsed —
// show" affordance that flips ?broadcast=show. 0 when none / ?broadcast=show.
BroadcastCollapsed int `json:"broadcastCollapsed,omitempty"`
}
ServiceTopologyResponse is the JSON shape served by /api/topology/service — flat node + edge lists matching the chstore ServiceTopologyEdge model. Includes the time window so the UI can render the active window and a draw.io export can embed it in the filename.
type SpanMetricServiceRow ¶ added in v0.5.350
type SpanMetricServiceRow struct {
Service string `json:"service"`
Calls uint64 `json:"calls"`
Errors uint64 `json:"errors"`
ErrorRate float64 `json:"errorRate"`
AvgMs float64 `json:"avgMs,omitempty"`
MaxMs float64 `json:"maxMs,omitempty"`
P50Ms float64 `json:"p50Ms,omitempty"`
P99Ms float64 `json:"p99Ms,omitempty"`
// Inline call-rate sparkline — 30 buckets evenly spread
// across the requested window. Float so a SVG renderer can
// scale by max() without integer truncation; counts are
// integers in the source data but we already pay the
// float math in the aggregation step.
Sparkline []float64 `json:"sparkline,omitempty"`
// Source metric names this row aggregated — surfaced to
// the UI so the operator can confirm which spanmetrics
// processor variant their collector is emitting.
CallsMetric string `json:"callsMetric,omitempty"`
DurationMetric string `json:"durationMetric,omitempty"`
}
SpanMetricServiceRow is one row of the span-metric-derived service overview — call volume + error fraction in the window, optionally augmented with histogram-derived latency once we wire that. Returned by /api/spanmetrics/services.
type TopologyNode ¶ added in v0.5.100
type TopologyNode struct {
ID string `json:"id"`
Service string `json:"service"`
Op string `json:"op"`
}
TopologyNode is a service.operation node in the response. The frontend keys nodes by `id` (service|op) so render-time edge lookups stay O(1).
type TopologyResponse ¶ added in v0.5.100
type TopologyResponse struct {
Nodes []TopologyNode `json:"nodes"`
Edges []chstore.TopologyEdge `json:"edges"`
RootService string `json:"rootService"`
Depth int `json:"depth"`
From int64 `json:"from"` // unix ns
To int64 `json:"to"`
Truncated bool `json:"truncated"`
}
TopologyResponse is the JSON shape served by /api/topology. Truncated is true when the underlying edge query hit the LIMIT — the UI shows a banner so the operator knows the view is partial.
Source Files
¶
- ai_feedback.go
- ai_observability.go
- alert_tuning.go
- annotation_routes.go
- announcement.go
- anomaly_extra.go
- anomaly_window.go
- api.go
- api_databases.go
- api_incidents.go
- api_logs.go
- api_monitors.go
- api_slo.go
- api_tokens.go
- cache.go
- ch_ddl_queue.go
- clickhouse_coordinators.go
- clickhouse_health.go
- clickhouse_nodework.go
- cluster.go
- config_iox.go
- copilot_aianalyze.go
- copilot_ch_optimize.go
- copilot_chat.go
- copilot_deps.go
- copilot_drawer.go
- copilot_exception.go
- copilot_followup.go
- copilot_guided.go
- copilot_nl_query.go
- copilot_schemas.go
- copilot_shift.go
- correlate.go
- correlation_link.go
- custom_roles.go
- dashboard_merge.go
- db_waitlock.go
- dbstmt_detail.go
- deploys_page.go
- dql.go
- endpoints_detail.go
- entity_scan.go
- evaluator_health.go
- events.go
- exception_burst.go
- explain_spans_pick.go
- explain_trace_input.go
- external.go
- greeting.go
- hosts.go
- inbox.go
- jsonsafe.go
- kibana_handlers.go
- kibana_sync.go
- ldap_groupsync.go
- logstore_es_handlers.go
- logstream_broker.go
- mcp_gate.go
- me_cache.go
- metricresolve.go
- notifications.go
- pages.go
- pipeline.go
- pivot.go
- podservice.go
- presence.go
- problem_assign_notify.go
- problem_counts_cache.go
- problem_enrich.go
- problems_filter.go
- promql.go
- purge.go
- rag.go
- rca_evidence.go
- rca_impact.go
- rca_record.go
- rca_shields.go
- rca_verdict.go
- request_id_links.go
- rollup_routes.go
- rootcause.go
- runbook.go
- runbook_exec.go
- service_metric_throughput.go
- service_runtimes_cache.go
- servicegraph.go
- shapes.go
- slo_autocreate.go
- sql_playground.go
- team_aliases.go
- tempo.go
- tempo_handlers.go
- thanos_handlers.go
- topology.go
- trace_resolve.go
- traces_extras.go
- watchers.go