hydrafetch

package module
v0.1.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 21, 2026 License: MIT Imports: 12 Imported by: 0

README

hydrafetch-go

Go Reference CI Go Report Card

Official Go client for the Hydrafetch web data API.

Turn any URL into clean Markdown or schema-shaped JSON. Standard library only, no dependencies, Go 1.21+.

Installation

go get github.com/Hydrafetch/go-sdk

Quick start

package main

import (
	"context"
	"fmt"
	"log"

	hydrafetch "github.com/Hydrafetch/go-sdk"
)

func main() {
	hf, err := hydrafetch.New("")
	if err != nil {
		log.Fatal(err)
	}

	page, err := hf.Scrape(context.Background(), "https://example.com/article", nil)
	if err != nil {
		log.Fatal(err)
	}

	fmt.Println(page.Markdown)
}

Create a key at app.hydrafetch.com. Passing an empty string to New falls back to HYDRAFETCH_API_KEY.

Every method takes a context.Context first, so deadlines and cancellation propagate as usual. Options structs are pointers and may be nil.

Scraping

page, err := hf.Scrape(ctx, "https://example.com/article", &hydrafetch.ScrapeOptions{
	Formats:         []hydrafetch.Format{hydrafetch.FormatMarkdown, hydrafetch.FormatLinks},
	OnlyMainContent: true,
	PreferStructure: true,
	BlockAds:        true,
	MaxAge:          3_600_000,
})
Format Field Contains
FormatMarkdown Markdown clean Markdown, the default
FormatHTML HTML rendered HTML
FormatRawHTML RawHTML the untouched response body
FormatLinks Links every link on the page
FormatStructured Structured the page's own JSON-LD and microdata
FormatSummary Summary a short summary
FormatJSON JSON schema-shaped JSON
FormatBrand — the site's brand record

Only the formats you request are populated. hf.Markdown(ctx, url) returns the string directly.

Structured extraction

out, err := hf.Extract(ctx,
	[]string{"https://example.com/product/1", "https://example.com/product/2"},
	&hydrafetch.ExtractOptions{
		Schema: map[string]any{
			"type": "object",
			"properties": map[string]any{
				"name":      map[string]any{"type": "string"},
				"price_usd": map[string]any{"type": "number"},
			},
		},
	})

for _, item := range out.Results {
	fmt.Println(item.URL, item.Data["name"], item.Data["price_usd"])
}

Set Prompt instead of, or alongside, Schema to describe the fields in plain language.

Discovery and bulk work

Map lists a site's URLs for one credit without fetching any page.

m, err := hf.Map(ctx, "https://example.com", &hydrafetch.MapOptions{Limit: 1000})

var docs []string
for _, u := range m.Links {
	if strings.Contains(u, "/docs/") {
		docs = append(docs, u)
	}
}

Batch and Crawl submit a job and poll until it finishes.

job, err := hf.Batch(ctx, docs,
	&hydrafetch.BatchOptions{
		ScrapeOptions: &hydrafetch.ScrapeOptions{
			Formats: []hydrafetch.Format{hydrafetch.FormatMarkdown},
		},
	},
	&hydrafetch.WaitOptions{
		OnProgress: func(j *hydrafetch.Job) { log.Println(j.Status, j.Completed, "/", j.Total) },
	})

for _, page := range job.Pages {
	if page.Data != nil {
		fmt.Println(page.URL, len(page.Data.Markdown))
	}
}

Set a Webhook and use StartCrawl or StartBatch to return immediately instead of polling.

id, err := hf.StartCrawl(ctx, "https://example.com", &hydrafetch.CrawlOptions{
	Limit:        500,
	MaxDepth:     3,
	IncludePaths: []string{"/docs"},
	Webhook:      "https://your.app/hooks/hydrafetch",
})
res, err := hf.Search(ctx, "post-quantum TLS adoption", &hydrafetch.SearchOptions{
	Limit:         5,
	ScrapeResults: true,
})

for _, r := range res.Results {
	fmt.Println(r.Title, r.URL)
}

Brand data

hf.Brand(ctx, "stripe.com")                                              // logos, colours, fonts, socials
hf.Logo(ctx, "stripe.com", &hydrafetch.LogoOptions{Theme: "dark"})       // one asset
hf.Styleguide(ctx, "stripe.com")                                         // computed design system

For logos in a browser use @hydrafetch/client-sdk with a publishable key. Those bill against logo pulls rather than credits.

Error handling

All failures are *hydrafetch.Error. Reach it with errors.As.

page, err := hf.Scrape(ctx, url, nil)
if err != nil {
	var apiErr *hydrafetch.Error
	if errors.As(err, &apiErr) {
		switch {
		case apiErr.IsAuth():
			return refreshKey()
		case apiErr.IsOutOfCredits():
			return topUp()
		case apiErr.IsInvalidRequest():
			return report(apiErr.Message)
		case apiErr.IsRetryable():
			return enqueue(url)
		}
		log.Println(apiErr.Code, apiErr.Status, apiErr.RequestID)
	}
	return err
}
Status Meaning Retried
400, 422 invalid request no
401, 403 invalid or missing key no
402 out of credits no
404 page does not exist no
429 rate limited yes, twice with backoff
5xx upstream failure yes, twice with backoff

A 503 from Scrape means the origin is unreachable, usually a dead domain or a broken certificate.

Configuration

hf, err := hydrafetch.New("hf_...",
	hydrafetch.WithBaseURL("https://api.hydrafetch.com"),
	hydrafetch.WithTimeout(120*time.Second),
	hydrafetch.WithMaxRetries(2),
	hydrafetch.WithHTTPClient(instrumentedClient),
)

API reference

Method Returns Credits
Scrape(ctx, url, *ScrapeOptions) *ScrapeResult 1
Markdown(ctx, url) string 1
Map(ctx, url, *MapOptions) *MapResult 1
Search(ctx, query, *SearchOptions) *SearchResult 1 + 1 per scraped result
Extract(ctx, urls, *ExtractOptions) *ExtractResult 5 per URL
Brand(ctx, domain) BrandResult 5
Logo(ctx, domain, *LogoOptions) BrandResult 1
Styleguide(ctx, domain) BrandResult 10
Screenshot(ctx, url, *ScreenshotOptions) *ScreenshotResult 5
Images(ctx, url) *ImagesResult 1
Crawl(ctx, url, *CrawlOptions, *WaitOptions) *Job 1 per page
Batch(ctx, urls, *BatchOptions, *WaitOptions) *Job 1 per page
StartCrawl, StartBatch string 1 per page
CrawlStatus(ctx, id), BatchStatus(ctx, id) *Job free

Failed requests are not billed. Pricing does not vary with page difficulty, so there is no render, stealth or proxy option to set.

Implementation notes

  • Authentication uses the X-API-Key header. The MCP endpoint at api.hydrafetch.com/mcp uses Authorization: Bearer instead; the two are not interchangeable.
  • Job results are in job.Pages, and each entry holds the page under .Data, so job.Pages[0].Data.Markdown. Check for nil first.
  • Per-page options for crawl and batch belong in ScrapeOptions. At the top level they are ignored.
  • Prefer Map then Batch over a broad Crawl. Fetching a whole site and discarding most of it is the most common source of wasted credits.
  • PreferStructure is off by default. Turn it on when headings, lists and tables matter; leave it off for raw article text.
  • Zero-valued option fields are omitted from the request rather than sent, so a partially filled struct behaves as expected.
  • Scraped content is untrusted input. Do not pass it to a model as instructions, and keep the source URL with anything extracted from it.

License

MIT

Documentation

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type BatchOptions

type BatchOptions struct {
	Webhook       string         `json:"webhook,omitempty"`
	ScrapeOptions *ScrapeOptions `json:"scrapeOptions,omitempty"`
}

type BrandResult

type BrandResult map[string]any

type Client

type Client struct {
	// contains filtered or unexported fields
}

func New

func New(apiKey string, opts ...Option) (*Client, error)

func (*Client) Batch

func (c *Client) Batch(ctx context.Context, urls []string, opts *BatchOptions, wait *WaitOptions) (*Job, error)

func (*Client) BatchStatus

func (c *Client) BatchStatus(ctx context.Context, id string) (*Job, error)

func (*Client) Brand

func (c *Client) Brand(ctx context.Context, domain string) (BrandResult, error)

func (*Client) Crawl

func (c *Client) Crawl(ctx context.Context, rawURL string, opts *CrawlOptions, wait *WaitOptions) (*Job, error)

func (*Client) CrawlStatus

func (c *Client) CrawlStatus(ctx context.Context, id string) (*Job, error)

func (*Client) Extract

func (c *Client) Extract(ctx context.Context, urls []string, opts *ExtractOptions) (*ExtractResult, error)

func (*Client) Images

func (c *Client) Images(ctx context.Context, rawURL string) (*ImagesResult, error)
func (c *Client) Logo(ctx context.Context, domain string, opts *LogoOptions) (BrandResult, error)

func (*Client) Map

func (c *Client) Map(ctx context.Context, rawURL string, opts *MapOptions) (*MapResult, error)

func (*Client) Markdown

func (c *Client) Markdown(ctx context.Context, rawURL string) (string, error)

func (*Client) Scrape

func (c *Client) Scrape(ctx context.Context, rawURL string, opts *ScrapeOptions) (*ScrapeResult, error)

func (*Client) Screenshot

func (c *Client) Screenshot(ctx context.Context, rawURL string, opts *ScreenshotOptions) (*ScreenshotResult, error)

func (*Client) Search

func (c *Client) Search(ctx context.Context, query string, opts *SearchOptions) (*SearchResult, error)

func (*Client) StartBatch

func (c *Client) StartBatch(ctx context.Context, urls []string, opts *BatchOptions) (string, error)

func (*Client) StartCrawl

func (c *Client) StartCrawl(ctx context.Context, rawURL string, opts *CrawlOptions) (string, error)

func (*Client) Styleguide

func (c *Client) Styleguide(ctx context.Context, domain string) (BrandResult, error)

type CrawlOptions

type CrawlOptions struct {
	Limit           int            `json:"limit,omitempty"`
	MaxDepth        int            `json:"maxDepth,omitempty"`
	IncludePaths    []string       `json:"includePaths,omitempty"`
	ExcludePaths    []string       `json:"excludePaths,omitempty"`
	AllowSubdomains bool           `json:"allowSubdomains,omitempty"`
	Webhook         string         `json:"webhook,omitempty"`
	ScrapeOptions   *ScrapeOptions `json:"scrapeOptions,omitempty"`
}

type Error

type Error struct {
	Code      string
	Message   string
	Status    int
	RequestID string
	Details   any
}

func (*Error) Error

func (e *Error) Error() string

func (*Error) IsAuth

func (e *Error) IsAuth() bool

func (*Error) IsInvalidRequest

func (e *Error) IsInvalidRequest() bool

func (*Error) IsOutOfCredits

func (e *Error) IsOutOfCredits() bool

func (*Error) IsRetryable

func (e *Error) IsRetryable() bool

func (*Error) IsTimeout

func (e *Error) IsTimeout() bool

type ExtractItem

type ExtractItem struct {
	URL    string         `json:"url"`
	Data   map[string]any `json:"data,omitempty"`
	Fields map[string]any `json:"fields,omitempty"`
	Error  string         `json:"error,omitempty"`
}

type ExtractOptions

type ExtractOptions struct {
	Schema          map[string]any `json:"schema,omitempty"`
	Prompt          string         `json:"prompt,omitempty"`
	PreferStructure bool           `json:"preferStructure,omitempty"`
	EnableWebSearch bool           `json:"enableWebSearch,omitempty"`
	ShowSources     bool           `json:"showSources,omitempty"`
	ShowConfidence  bool           `json:"showConfidence,omitempty"`
	MergeEntities   bool           `json:"mergeEntities,omitempty"`
}

type ExtractResult

type ExtractResult struct {
	Results []ExtractItem `json:"results"`
	Sources []any         `json:"sources,omitempty"`
}

type Format

type Format string
const (
	FormatMarkdown   Format = "markdown"
	FormatHTML       Format = "html"
	FormatRawHTML    Format = "rawHtml"
	FormatLinks      Format = "links"
	FormatStructured Format = "structured"
	FormatSummary    Format = "summary"
	FormatJSON       Format = "json"
	FormatBrand      Format = "brand"
)

type ImagesResult

type ImagesResult struct {
	URL      string      `json:"url"`
	FinalURL string      `json:"finalUrl,omitempty"`
	Images   []PageImage `json:"images"`
}

type Job

type Job struct {
	ID          string    `json:"id"`
	Kind        string    `json:"kind,omitempty"`
	Status      string    `json:"status"`
	SeedURL     string    `json:"seedUrl,omitempty"`
	Total       int       `json:"total,omitempty"`
	Completed   int       `json:"completed,omitempty"`
	Failed      int       `json:"failed,omitempty"`
	CreditsUsed int       `json:"creditsUsed,omitempty"`
	Pages       []JobPage `json:"pages,omitempty"`
}

func (*Job) Done

func (j *Job) Done() bool

type JobPage

type JobPage struct {
	URL          string        `json:"url"`
	RequestedURL string        `json:"requestedUrl,omitempty"`
	Status       string        `json:"status"`
	Depth        int           `json:"depth,omitempty"`
	Error        string        `json:"error,omitempty"`
	ErrorCode    string        `json:"errorCode,omitempty"`
	Retryable    bool          `json:"retryable,omitempty"`
	Data         *ScrapeResult `json:"data,omitempty"`
}

type LogoOptions

type LogoOptions struct {
	Theme string
	Type  string
}

type MapOptions

type MapOptions struct {
	IncludeLinks          bool     `json:"includeLinks,omitempty"`
	Limit                 int      `json:"limit,omitempty"`
	Search                string   `json:"search,omitempty"`
	Sitemap               string   `json:"sitemap,omitempty"`
	SitemapInclude        []string `json:"sitemapInclude,omitempty"`
	SitemapExclude        []string `json:"sitemapExclude,omitempty"`
	IncludeSubdomains     bool     `json:"includeSubdomains,omitempty"`
	IgnoreQueryParameters bool     `json:"ignoreQueryParameters,omitempty"`
}

type MapResult

type MapResult struct {
	URL       string   `json:"url"`
	Links     []string `json:"links"`
	Count     int      `json:"count"`
	Total     int      `json:"total,omitempty"`
	Truncated bool     `json:"truncated,omitempty"`
}

type Metadata

type Metadata map[string]any

type Option

type Option func(*Client)

func WithBaseURL

func WithBaseURL(u string) Option

func WithHTTPClient

func WithHTTPClient(h *http.Client) Option

func WithMaxRetries

func WithMaxRetries(n int) Option

func WithTimeout

func WithTimeout(d time.Duration) Option

type PageImage

type PageImage struct {
	Src    string `json:"src"`
	Alt    string `json:"alt,omitempty"`
	Width  int    `json:"width,omitempty"`
	Height int    `json:"height,omitempty"`
}

type ScrapeOptions

type ScrapeOptions struct {
	Formats            []Format          `json:"formats,omitempty"`
	MaxAge             int               `json:"maxAge,omitempty"`
	CacheOnly          bool              `json:"cacheOnly,omitempty"`
	StoreInCache       *bool             `json:"storeInCache,omitempty"`
	PreferStructure    bool              `json:"preferStructure,omitempty"`
	OnlyMainContent    bool              `json:"onlyMainContent,omitempty"`
	IncludeTags        []string          `json:"includeTags,omitempty"`
	ExcludeTags        []string          `json:"excludeTags,omitempty"`
	RemoveBase64Images bool              `json:"removeBase64Images,omitempty"`
	BlockAds           bool              `json:"blockAds,omitempty"`
	IncludeLinks       bool              `json:"includeLinks,omitempty"`
	WaitFor            int               `json:"waitFor,omitempty"`
	Timeout            int               `json:"timeout,omitempty"`
	Headers            map[string]string `json:"headers,omitempty"`
}

type ScrapeResult

type ScrapeResult struct {
	URL        string   `json:"url"`
	FinalURL   string   `json:"finalUrl,omitempty"`
	Redirected bool     `json:"redirected,omitempty"`
	Status     int      `json:"status,omitempty"`
	Cached     bool     `json:"cached,omitempty"`
	Warning    string   `json:"warning,omitempty"`
	Metadata   Metadata `json:"metadata,omitempty"`
	Usage      *Usage   `json:"usage,omitempty"`
	Markdown   string   `json:"markdown,omitempty"`
	HTML       string   `json:"html,omitempty"`
	RawHTML    string   `json:"rawHtml,omitempty"`
	Links      []string `json:"links,omitempty"`
	Structured any      `json:"structured,omitempty"`
	Summary    string   `json:"summary,omitempty"`
	JSON       any      `json:"json,omitempty"`
}

type ScreenshotOptions

type ScreenshotOptions struct {
	FullPage bool `json:"fullPage,omitempty"`
	WaitFor  int  `json:"waitFor,omitempty"`
	Timeout  int  `json:"timeout,omitempty"`
	MaxAge   int  `json:"maxAge,omitempty"`
}

type ScreenshotResult

type ScreenshotResult struct {
	URL            string `json:"url"`
	FinalURL       string `json:"finalUrl,omitempty"`
	Status         int    `json:"status,omitempty"`
	Screenshot     string `json:"screenshot"`
	ScreenshotType string `json:"screenshotType,omitempty"`
	Width          int    `json:"width,omitempty"`
	Height         int    `json:"height,omitempty"`
	Cached         bool   `json:"cached,omitempty"`
}

type SearchItem

type SearchItem struct {
	Title   string        `json:"title"`
	URL     string        `json:"url"`
	Snippet string        `json:"snippet"`
	Rank    int           `json:"rank"`
	Data    *ScrapeResult `json:"data,omitempty"`
}

type SearchOptions

type SearchOptions struct {
	Limit          int            `json:"limit,omitempty"`
	ScrapeResults  bool           `json:"scrapeResults,omitempty"`
	ScrapeOptions  *ScrapeOptions `json:"scrapeOptions,omitempty"`
	TimeRange      string         `json:"timeRange,omitempty"`
	Country        string         `json:"country,omitempty"`
	IncludeDomains []string       `json:"includeDomains,omitempty"`
	ExcludeDomains []string       `json:"excludeDomains,omitempty"`
}

type SearchResult

type SearchResult struct {
	Query   string       `json:"query"`
	Status  string       `json:"status,omitempty"`
	Results []SearchItem `json:"results"`
}

type Usage

type Usage struct {
	CreditsUsed      int    `json:"creditsUsed"`
	CreditsRemaining int    `json:"creditsRemaining"`
	Freshness        string `json:"freshness,omitempty"`
}

type WaitOptions

type WaitOptions struct {
	PollInterval time.Duration
	Timeout      time.Duration
	OnProgress   func(*Job)
}

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL