mipstack

package module
v0.0.0-...-961d4b1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 30, 2026 License: MPL-2.0 Imports: 20 Imported by: 5

README

mihomo IP stack

The mihomo IP stack (MIPS) is the small, pure-Go userspace IP stack developed for mihomo. It converts complete IPv4 and IPv6 packets from an arbitrary L3 link into standard Go TCP, UDP, and IP protocol socket interfaces. It requires Go 1.20 or later, uses only the Go standard library, and does not require cgo.

The package is independent of any particular link, routing policy, or L2 neighbor implementation. An embedding application owns route admission and packet delivery.

Usage

package main

import (
    "context"
    "net/netip"

    "github.com/metacubex/mipstack"
)

func open(ctx context.Context, local netip.Prefix, destination netip.AddrPort) error {
    stack, err := mipstack.New(mipstack.Config{
        LocalAddresses: []netip.Prefix{local},
        MTU:            1500,
        TCP: mipstack.TCPSocketDefaults{
            CongestionControl: mipstack.CongestionControlCUBIC,
        },
    })
    if err != nil {
        return err
    }
    defer stack.Close()
    if err = stack.Start(); err != nil {
        return err
    }

    connection, err := stack.DialTCP(ctx, "tcp", netip.AddrPort{}, destination)
    if err != nil {
        return err
    }
    return connection.Close()
}

Stack.Write delivers complete inbound IP packets to the stack. Stack.Read returns complete outbound IP packets for the embedding link:

_, err := stack.Write([][]byte{inboundPacket}, 0)

buffer := make([]byte, 65535)
sizes := []int{0}
_, err = stack.Read([][]byte{buffer}, sizes, 0)
outboundPacket := buffer[:sizes[0]]

The packet methods intentionally match the common batched userspace-TUN shape. Read blocks for one packet and then drains up to 64 currently queued packets into the supplied buffers; BatchSize reports that upper bound. Write accepts every packet supplied in an inbound batch. Start is idempotent, while Close is terminal and unblocks pending packet and socket operations. It returns after publishing closure and starting orderly shutdown; it does not wait for TCP actors or user forwarder handlers to return. Stack-owned queues and caches are discarded before it returns, while actor-owned buffers are released as those actors observe cancellation. Read requires sizes to be at least as long as the buffer slice and honors the same leading offset in every buffer. Only sizes entries below the returned count are valid; later entries are unchanged. Both methods report the successfully completed packet prefix before any later-buffer error, as expected by wireguard-go's packet-device loops. A short Read destination consumes and discards that packet before returning io.ErrShortBuffer. Invalid, unrelated, and unsupported packets passed to Write are accounted as drops but count as successfully consumed. Write accepts a buffer slice larger than BatchSize because a composite WireGuard device may use a larger Bind batch; Read also accepts such a slice but still returns no more than 64 packets.

Read and Write may run concurrently, including multiple calls to the same method. Concurrent calls have no relative completion or processing order, and each queued outbound packet is assigned to at most one Read. The Stack does not retain input buffers after Write returns or access destination buffers after Read returns, so callers may immediately reuse them.

The outbound link queue uses byte-based deficit round robin modeled on the local-flow scheduling in Linux sch_fq. New flows receive a bounded initial allowance, active flows rotate by byte credit, and short UDP or ICMP exchanges therefore do not sit behind an entire queue of bulk TCP packets. TCP pacing remains connection-owned: the scheduler only chooses among packets that a TCP actor has already made eligible. Local loopback delivery remains FIFO because it has no serialized external-link bottleneck.

When the fixed link queue is full of published packets, flow-aware admission prevents an earlier bulk flow from excluding later UDP, IP, and control flows from the scheduler.

The link queue is finite. If Stack.Read stops, TCP retains protocol output until capacity returns while its actors remain responsive and stream writes continue to obey their send-buffer and deadline rules. UDP and IP socket writes instead make one immediate bounded admission attempt. Failure to admit unicast output or an external-link non-unicast copy is silent by default and reports ENOBUFS when ReceiveErrors is enabled. Receive-side multicast and broadcast loopback copies remain independently best effort. Best-effort control packets may displace queued backlog or be discarded, and Stack.Write itself does not wait for outbound capacity. A successful socket write accepts the message but does not guarantee that every resulting packet reaches Stack.Read.

SetRXChecksumOffload delegates selected IPv4 header, TCP, UDP, ICMP, or IGMP checksum verification to a trusted input link; RXChecksumOffload returns the current policy. The zero value retains software checksum verification. Configure offload before delivering input and only when the link supplies valid complete packets or an equivalent checksum guarantee. Framing checks, IPv6 UDP's nonzero checksum requirement, reassembled transport checksums, and public codec validation remain enabled.

For integration with userspace packet-device consumers, Stack also provides MTU, Name, and BatchSize. LocalAddresses returns an independent snapshot of every configured address in configuration order. Operating-system file descriptors and event channels whose element type belongs to another package are left to embedding adapters so MIPS remains standard-library-only.

Packet codec

ParseIPPacket exposes a validated, zero-copy IPPacket view of one complete wire IPv4 or IPv6 packet, including a single fragment. It does not perform stateful reassembly. Its TCPSegment, UDPDatagram, and ICMPMessage methods validate the final upper-layer protocol and checksum and return the matching semantic value only for an unfragmented or reassembled packet. Parsed option and payload slices borrow the input packet; callers must copy or replace a slice before changing input they do not own.

The same four values construct wire data through MarshalBinary and AppendBinary. They implement encoding.BinaryMarshaler on every supported Go version and encoding.BinaryAppender when built with Go 1.24 or newer; the AppendBinary method remains directly callable with Go 1.20. Both methods validate the complete value, calculate the checksum owned by that layer, and do not retain caller storage. IPPacket calculates the IPv4 header checksum but leaves upper-layer checksums to TCPSegment, UDPDatagram, or ICMPMessage; callers building another protocol can use the checksum helpers below. AppendBinary preserves the existing destination prefix and supports output that overlaps borrowed option or payload storage, including appending to a zero-length view of the parsed wire buffer. It returns the original destination unchanged when validation fails. MarshalBinary is semantically identical to AppendBinary(nil). For TCP, UDP, and ICMP, source and destination addresses provide address-family and pseudo-header checksum context but are not part of the returned transport wire. IPv4-mapped input addresses are normalized to IPv4 during construction, while an IPv4-mapped address encoded in an IPv6 header is rejected, including inside a quoted packet carried by an ICMP error. TCP encoding clears the three unexposed reserved bits and normalizes every byte after End of Option List to RFC 9293's required zero padding while parsing remains compatible with Linux's tolerant receive behavior. The historic NS bit remains explicitly available. IPv6 encoding similarly clears PadN data and Fragment reserved fields without hiding the received bytes from a parsed IPPacket.

IPPacket.MarshalRawBinary and AppendRawBinary instead encode a valid fixed IP header while treating IPv4 options and the IPv6 Protocol and Payload as opaque wire data. They preserve IPv4 End padding and IPv6 extension-header reserved fields, and allow malformed IPv4 option framing and representable but semantically invalid fragment payloads for protocol testing. The corresponding SetRawIPv6ExtensionHeaders links recognized extension descriptors without enforcing their framing, order, uniqueness, option, or Fragment semantics. These opt-in methods do not produce malformed fixed headers; callers testing such fields can mutate the owned result. They also do not perform automatic raw fragmentation: MarshalFragments remains the strict source-fragmentation planner, while callers can encode or mutate each deliberately invalid fragment.

IPv4 option parsing likewise follows Linux's tolerant EOL behavior: received bytes after End remain available in IPPacket.IPv4Options, while structured option traversal stops at End. Strict packet encoding writes canonical zero padding, while raw encoding preserves the complete supplied option area.

ICMPMessage.IsEchoRequest and IsEchoReply identify complete IPv4 and IPv6 Echo messages, while Echo returns their identifier, sequence, and a borrowed payload view. SetEchoRequest and SetEchoReply build either direction with an independently owned payload. EchoReply creates a zero-copy semantic reply from an existing request using a caller-selected source address; explicit selection is required because multicast, broadcast, and anycast destinations cannot be reused blindly as reply sources. The zero-copy reply shares Body with the request, while MarshalBinary and AppendBinary encode the complete message and calculate its address-family checksum. ICMPMessage.IsError classifies supported family-specific error type/code pairs, while ICMPMessage.ICMPError validates the available quoted structure and returns its addresses, protocol, TCP or UDP ports, path MTU, and parameter pointer. Quoted packet and payload slices borrow ICMPMessage.Body; socket delivery takes an independent copy when it must retain them. For IPv6 No Next Header, QuotedPacket retains ignored trailing bytes while QuotedPayload is empty, matching IPPacket.UpperLayer. ICMPError.ICMPMessage performs the reverse construction from a validated, possibly truncated quoted packet and copies that quote; route selection, rate limiting, recursive-error suppression, and quote truncation remain transmission policy rather than codec behavior. The exported untyped ICMP type and code constants cover every error subtype the stack accepts as well as Echo Request and Reply. This includes the RFC 8883 IPv6 Parameter Problem processing-limit codes and Destination Unreachable "Headers too long", plus RFC 9914 P-Route errors. ICMPError.Extensions retains the RFC 4884 object sequence without its four-byte Extension Header; ExtensionObjects exposes ordered borrowed ICMPExtensionObject views and SetExtensionObjects performs the copying reverse conversion. Unknown and repeated objects remain lossless. ICMPExtensionObject.Pointer and SetPointer handle RFC 8883's Extended Information Pointer object, including pointers beyond the available quote. Encoding supplies the version, checksum, 128-byte minimum quotation, and family-specific four- or eight-byte padding; decoding verifies those fields and never guesses an extension when the RFC 4884 Length field is zero.

TCP options are available in wire order through TCPSegment.HeaderOptions and SetHeaderOptions. TCPHeaderOption preserves unknown and repeated kinds and provides typed construction and inspection for MSS, Window Scale, SACK-Permitted, SACK blocks, and Timestamps. TCPSACKBlock uses the original wrapping 32-bit sequence edges without applying connection-specific window policy. IPv4 options use the corresponding IPv4HeaderOptions and SetIPv4HeaderOptions methods; IPv4HeaderOption also exposes the copied, class, and number fields of its complete option type. IPv4 and IPv6 option descriptors both provide typed RouterAlert and SetRouterAlert methods; the standalone codec preserves the complete 16-bit value, while stack IGMP and MLD handling recognizes the protocol-defined value zero. ProtocolIGMP exposes the corresponding IP protocol number alongside the existing protocol constants.

IPPacket.IPv6ExtensionHeaders and SetIPv6ExtensionHeaders expose and build the linked extension-header sequence without making callers write Next Header links themselves. Hop-by-Hop and Destination Options headers additionally use IPv6ExtensionHeader.Options and SetOptions for their option TLVs. Set methods copy caller data and leave their receiver unchanged on failure; parse methods return caller-owned descriptors whose data slices borrow the original wire storage. Parsing preserves received sender-reserved fields for inspection; serialization clears PadN data and Fragment reserved fields. Authentication and Mobility headers remain opaque: their reserved fields are integrity-protected, so callers constructing either header must provide canonical fields together with a matching ICV or checksum.

IPPacket.Fragment returns a borrowed IPPacketFragmentView with byte-based offset, common 32-bit identification, More Fragments state, the fragment's Next Header value, and its raw fragmentable payload. IPv4 exposes the header fields directly on IPPacket; IPv6ExtensionHeader.Fragment and SetFragment provide typed access to the eight-byte IPv6 header while preserving the structural API's rule that Data excludes Next Header. An IPv6 atomic Fragment remains visible through this view and IsAtomic, but may still be traversed by UpperLayer. For a non-atomic IPv6 fragment, IPv6ExtensionHeaders returns the Fragment header as its final descriptor, its Next Header value as protocol, and the remaining fragment bytes without interpreting them as another complete extension header. A packet contains at most one structurally visible Fragment header. Standalone validation bounds the fragmentable-data end independently of the current fragment's header length; only stateful reassembly can apply the header retained from the offset-zero fragment, because RFC 8200 permits those headers to differ between fragments.

IPPacket.MarshalFragments performs stateless source fragmentation against an explicit L3 MTU and returns an all-or-nothing set of owned wire packets. A fitting input is identical to MarshalBinary. IPv4 uses the packet's existing Identification, honors DF, preserves copied options using Linux-compatible NOP replacement, and supports RFC 791 refragmentation by accumulating the original offset and MF state. IPv6 inserts the Fragment header after the RFC 8200 Per-Fragment headers, uses the caller-supplied identification, and replaces an existing atomic header rather than nesting one. For a newly fragmented datagram, the caller must provide an IPv4 Identification suitable for the source/destination/protocol tuple or an IPv6 identification suitable for the source and final destination, including the applicable non-reuse requirement. It refuses to refragment a non-atomic IPv6 fragment. The first IPv6 fragment contains every traversable extension header and the complete known upper-layer header as required by RFC 7112; an MTU too small for that chain returns syscall.EMSGSIZE. Known-size upper-layer headers comprise TCP, UDP, ICMP, IGMP, ESP, DCCP, SCTP, UDP-Lite, and nested IPv4 or IPv6. An unknown raw protocol remains fragmentable because its header boundary cannot be inferred from its protocol number alone.

IPPacketReassembly incrementally reconstructs one non-atomic fragmented packet. Its zero value is ready for use, and Add copies every retained byte so the caller may immediately reuse storage borrowed by the input IPPacket. Fragments may arrive out of order. A range already fully covered by retained fragments is treated as a duplicate without comparing its payload bytes or using its ECN and header metadata. Following Linux, a duplicate final fragment may establish the packet's final length, but the duplicate itself never completes reassembly. A partial overlap or other associated conflict invalidates the entire in-progress reassembly. Invalid standalone input and a fragment belonging to a different packet leave existing state unchanged. A successful result owns all option and payload storage, contains no non-atomic fragmentation state, and resets the receiver for reuse. Reset explicitly abandons an incomplete packet. The type deliberately has no clock, goroutine, global capacity policy, or Close: callers multiplexing many datagrams remain responsible for their identity table, lifetime, and aggregate resource limits. Stack supplies that surrounding policy for live network traffic.

IPPacket.UpperLayer structurally walks IPv6 Hop-by-Hop, Destination Options, Routing, atomic Fragment, Authentication, and Mobility headers. ESP and unknown values terminate traversal, while No Next Header returns no upper-layer payload even if preserved trailing bytes are available through the structural extension API. Structural AH or Mobility traversal does not authenticate or otherwise implement those protocols; AH framing still requires its complete fixed fields and IPv6's eight-octet alignment. IPv6 jumbograms are not supported. Non-atomic fragments remain structurally parseable and round-trip encodable, but UpperLayer and the TCP, UDP, and ICMP decoders reject them until IPPacketReassembly or Stack supplies a complete packet. Every Jumbo Payload option is rejected. A decoder that must validate an address-dependent pseudo-header (TCP, checksummed UDP, or ICMPv6) rejects an active IPv4 source route, active IPv6 Routing Header, or Mobile IPv6 Home Address option because it lacks the corresponding routing state. ICMPv4 and checksum-disabled IPv4 UDP remain decodable.

Protocol numbers, TCP flags, TCP and IP option kinds, and IPv6 extension-header identifiers are exposed as untyped constants so callers can use them directly with the wire-sized integer fields. This includes ProtocolESP and ProtocolNoNextHeader. TCPSegment.Flags remains a uint16. The package also exposes InternetChecksum and IPTransportChecksum for callers building other upper-layer protocols. Their InternetChecksumParts and IPTransportChecksumParts counterparts checksum logically contiguous input without first gathering its parts; empty parts preserve byte alignment across part boundaries. The codec intentionally stops at the 65,535-byte non-jumbogram IP model used by the Stack.

Socket API

MIPS provides:

  • DialTCP for active IPv4 and IPv6 TCP connections;
  • ListenTCP for specific or wildcard passive TCP endpoints;
  • DialUDP for connected UDP sockets;
  • ListenUDP for unconnected UDP packet sockets;
  • ListenMulticastUDP for a reusable wildcard UDP binding joined to one IPv4 or IPv6 multicast group;
  • DialIP and ListenIP for connected and unconnected IPv4 or IPv6 protocol payload sockets, using standard network names such as ip4:icmp, ip6:ipv6-icmp, and ip:99;
  • TCPForwarder, UDPForwarder, IPForwarder, and ICMPForwarder for otherwise unhandled inbound traffic, including transparent nonlocal destinations;
  • exported TCPConn, TCPListener, UDPConn, and IPConn implementations of the corresponding standard net interfaces.

The Stack dial methods and their Dialer counterparts return net.Conn. Stack.ListenTCP and ListenConfig.ListenTCP return net.Listener, while ordinary UDP and IP listen methods return net.PacketConn. Their dynamic types are *TCPConn, *TCPListener, *UDPConn, and *IPConn as appropriate, and TCPListener.Accept returns a net.Conn with dynamic type *TCPConn. Callers only need a type assertion when using MIPS-specific extensions. ListenMulticastUDP returns *UDPConn.

TCPConn.SetQuickACK mirrors Linux's transient TCP_QUICKACK policy. Enabling it replenishes a bounded immediate-ACK budget, leaves response-piggybacking mode, and queues an immediate flush of any pending acknowledgement; disabling it favors response piggybacking. Protocol events and the delayed-ACK timer may subsequently change the mode, so the setting is not persistent. A successful call queues the actor-owned change without waiting for the acknowledgement to reach the embedding link.

The zero-value ListenConfig and Dialer mirror the creation-time policy pattern used by net.ListenConfig and net.Dialer. Their Options slices are read in order and are not retained. SocketOptions constructs sealed, strongly typed policies without adding a top-level exported type for every option. The creation policies are:

Scope Options
TCP, UDP, and IP ReadBuffer, TrafficClass, FlowLabel
TCP connections WriteBuffer, KeepAlive, KeepAliveConfig, NoDelay, IdleTimeout, UserTimeout, CongestionControl, CongestionControlFactory, MaximumPacingRate
TCP listeners AcceptQueue, SYNBacklog, plus the TCP connection policies inherited by accepted connections
UDP and IP ReceiveErrors, PathMTUDiscovery, HopLimit, Broadcast, MulticastHopLimit, MulticastLoopback
TCP and UDP listeners ReuseAddress, ReusePort
IP IPHeaderIncludedOnWrite, IPHeaderIncludedOnRead, ICMPv4Filter, ICMPv6Filter, IPv6Checksum

Every constructor is an explicit choice, including boolean false and valid numeric zero values. Every policy has a corresponding UnsetXxx constructor; the congestion-control name and factory forms share UnsetCongestionControl. An unset marker removes an earlier choice of the same kind and restores the operation-specific Stack default. Unset markers are accepted by every socket creation operation and have no effect where the corresponding setting is inapplicable. This supports layered option slices without treating an explicit disable or zero as absence.

An option used with an inapplicable protocol or operation reports ENOPROTOOPT before an endpoint is created. Repeated option kinds use the last value or unset marker. IPConn.SetIPHeaderIncludedOnWrite may change the write representation of an existing socket between operations; SocketOption values are consumed only during socket creation. Like Linux raw sockets, an ICMP error delivered after such a change uses the write representation in effect when the error arrives. The direct Stack methods remain concise default-policy entry points.

A TCP listener retains only its explicit connection-policy overrides. Each accepted connection reads the current Config.TCP defaults at creation and then applies those overrides, so a later UpdateConfig affects future accepts without discarding listener-specific choices. Explicit TCP buffer capacities also become their auto-tuning maxima. The listener's AcceptQueue and SYNBacklog are fixed when it is bound, and its SYN-cookie responses use the same receive-window, Traffic Class, and Flow Label policy as stateful handshakes.

The three DialTCP, DialUDP, and DialIP entry points mirror the netip-based methods available on newer net.Dialer versions: they accept a context, network name, local address, and remote address, while returning net.Conn for straightforward adapter use. The package remains buildable with Go 1.20.

Listen methods accept the standard tcp, tcp4, tcp6, udp, udp4, and udp6 network names. An empty netip.Addr has the same wildcard meaning as a nil IP in net.TCPAddr or net.UDPAddr; its port is retained. For the generic network, a wildcard becomes one dual-stack [::] endpoint when both address families are configured. The 4 and 6 forms select one family and reject an explicit address from the other family.

ListenIP applies the same empty-address rules to ip, ip4, and ip6 protocol sockets. A generic ip:* wildcard receives both families when both are configured. On unconnected UDP and IP sockets, Read reads a payload and discards its source address, matching the corresponding standard connection types; ReadFrom and message reads retain it.

Direct TCP listeners enable ReuseAddress by default, matching Go's standard listener setup; direct UDP listeners use exclusive bindings. A ListenConfig can override the TCP default or opt UDP into reuse. Two overlapping UDP bindings are compatible only when both enable ReuseAddress or both enable ReusePort. A group in which every member enables ReusePort uses a per-registry keyed flow hash; a ReuseAddress group otherwise delivers unicast to its most recently bound member. TCP permits simultaneous shared listeners only through ReusePort. Exact bindings take precedence over wildcards. Port zero always allocates a distinct unused ephemeral port even when reuse is enabled. Accepted TCP connections inherit both policies, so a listener can be rebound over a live accepted connection only when the old and new sockets share the applicable Linux reuse policy.

UDP and raw IP sockets expose the single-interface equivalents of Go's x/net/ipv4 and x/net/ipv6 multicast controls: JoinGroup, LeaveGroup, the source-specific join/leave and include/exclude operations, and atomic SetMulticastSourceFilter snapshots for previously joined groups. MulticastSourceFilterExclude and MulticastSourceFilterInclude use Linux's MCAST_EXCLUDE=0 and MCAST_INCLUDE=1 values. No interface argument is needed because one Stack represents exactly one embedding interface.

Memberships belong to the socket that created them and are removed when that socket is closed. While the Stack remains operational, removing its final membership schedules the applicable state-change Report, Leave, or Done on a best-effort basis. Stack.Close is terminal: it cancels pending reports and packet I/O without attempting a final leave because the embedding link may no longer be usable. JoinSourceSpecificGroup creates an INCLUDE membership for its first source, and removing its final source leaves the group. The include/exclude delta operations require an existing membership in the corresponding mode. Closing one reusable listener does not affect memberships owned by the other listeners on that port.

As an RFC 4604 SSM-aware host, MIPS requires source-specific INCLUDE memberships for IPv4 232/8 and IPv6 FF3x::/32. Any-source joins, EXCLUDE filters, and EXCLUDE delta operations in those ranges return EINVAL; an older IGMPv1/v2 or MLDv1 Report cannot suppress a pending SSM report.

SetMulticastHopLimit and SetMulticastLoopback control output independently from ordinary unicast hop limits. ListenMulticastUDP disables loopback on the returned sending socket like net.ListenMulticastUDP; other sockets keep the standard enabled default. IPv4 limited and configured-subnet broadcasts are delivered to eligible wildcard UDP and raw sockets. SetBroadcast controls output permission and starts enabled, matching Go's default UDP and raw sockets on supported operating systems.

The stack maintains the RFC 9776 and RFC 3810 aggregate interface filter and emits IGMPv1/v2/v3 or MLDv1/v2 reports according to the active querier compatibility mode. Reports include the required Router Alert, fit the link MTU, and preserve source-filter state across socket and address changes. IPv6 interface-local multicast and an explicit multicast hop limit of zero remain inside the host; multicast loopback still obeys each sending socket's setting.

Transparent interception

Config.Promiscuous admits valid unicast IP packets whose destination is not listed by LocalAddresses. This is an L3 transparent-receive policy, not an L2 interface or MAC promiscuous mode. Local ownership remains unchanged: ordinary wildcard sockets receive only managed local destinations, source selection uses only LocalAddresses, and nonlocal destinations never become loopback routes. The forwarder name refers to delivery into an application handler; MIPS still does not route or forward IP packets between links.

Forwarders and promiscuous admission are independent. A forwarder always sees otherwise unhandled traffic addressed to LocalAddresses, without requiring Promiscuous. Promiscuous is required only when the original destination is not locally owned. Enabling it without a matching forwarder does not make ordinary wildcard sockets transparent and does not generate automatic replies; admitted nonlocal traffic without a matching protocol forwarder is silently dropped.

LocalAddresses may be empty when Promiscuous is enabled. Such a Stack has no addresses for ordinary socket binding, source selection, multicast, or loopback delivery; only forwarder-created endpoints and forwarder reply or reject actions may emit from intercepted addresses. Routes == nil installs source-less IPv4 and IPv6 default routes in this configuration. An explicit route slice can limit return destinations, and a non-nil empty slice permits interception but causes forwarder actions requiring a return path to report syscall.ENETUNREACH.

NewTCPForwarder, NewUDPForwarder, NewIPForwarder, and NewICMPForwarder install one protocol-specific fallback handler each. All four constructors take their own options type so later protocol policy can evolve independently. Established tuples and ordinary local listeners take precedence. A TCP request represents a valid initial SYN. Its handler runs in a new goroutine and may block, but an undecided handler occupies TCPForwarderOptions.MaxInFlight capacity. Accept blocks until the handshake, context cancellation, or stack closure and returns a TCPConn whose LocalAddr preserves the original destination. The handler must wait for Accept itself to return, but may return immediately afterward: the returned connection has its own lifetime and may instead be handed to another goroutine. Drop consumes the SYN silently, while Reject sends the RFC 9293 reset on a best-effort basis. Retransmitted SYNs do not create duplicate handler calls. Accept takes the same variadic connection SocketOption values as Dialer; option validation happens before the request is claimed, so an invalid option does not prevent a different terminal action.

TCPForwarderRequest.Done closes when its handler returns, a pending request is invalidated by configuration, or its forwarder or stack closes. Forwarder closure remains observable after the request has selected an action, allowing blocking handler work to stop promptly. Each forwarder's Done channel closes on direct closure or Stack.Close; Close does not wait for handlers, accepted endpoints, or replies already in progress.

UDP requests expose the source, original destination, and triggering payload. Accept creates a connected UDPConn, offers a copy of that first datagram to its receive queue, and retains the complete four-tuple for subsequent dispatch. The connection is bound to the original destination and connected to the original source: Read accepts only that peer, Write replies to it, and WriteTo reports net.ErrWriteToConnected. Listen instead creates an unconnected UDPConn bound to the original destination: it receives all later sources through ReadFrom and replies through WriteTo. Both actions offer the first datagram to the new socket's capacity-bounded receive queue; a queue drop does not undo endpoint registration. The handler must wait for Accept or Listen to return, but the returned connection remains valid after the callback. Both creation methods accept the UDP/IP policy SocketOption values used by Dialer; they validate options before claiming the request. Forwarded listeners retain exclusive flow ownership and therefore reject bind-reuse options. Reply writes one reverse datagram without retaining a flow and uses Flow().Destination as its default source. ReplyFrom selects any valid same-family source address and UDP port. Source membership in LocalAddresses and unicast, multicast, or broadcast classification are not policy checks; a zero source port is preserved. This is stateless transparent output, not arbitrary destination selection: every reply still targets Flow().Source. Replies may be repeated or retried and do not prevent a later Accept, Listen, Detach, DetachForReplies, Drop, or Reject. Stateless replies inherit the current Config.UDP output defaults.

ICMP requests similarly expose a checksum-validated complete message and provide a repeatable, checksum-repairing reverse Reply. IPPacket exposes the complete reassembled L3 packet: request-scoped storage is read-only and valid only during the callback, while a detached responder owns its mutable snapshot. In both forms Message().Payload aliases the ICMP region. Detached Source, Destination, Type, and Code remain the original metadata; changing the corresponding payload bytes makes IsEchoRequest false but changing IP header bytes does not reclassify that metadata. ICMPForwarderMessage.ICMPMessage validates the current wire snapshot and returns a zero-copy semantic view whose Body aliases Payload[4:]. SetICMPMessage performs the reverse conversion, recalculates the checksum, and reuses existing Payload capacity when possible; validation failure leaves the forwarder message unchanged.

ReplyIPPacket is a restricted header-included ICMP reply, not arbitrary raw packet injection. Its destination must be the triggering packet's source; its source may be any valid same-family address, without stack-membership or address classification policy. The stack copies caller storage, normalizes IPv4 Total Length or IPv6 Payload Length, repairs the outer IPv4 and ICMP checksum, and preserves other supported header fields. A fitting IPv6 atomic Fragment header is retained with reserved fields cleared. If the packet must be fragmented, an existing atomic header is replaced rather than nested, and the Fragment header is inserted after the RFC 8200 Per-Fragment header chain. IPv4 DF reports syscall.EMSGSIZE; otherwise source fragmentation preserves copied IPv4 options. ICMPv6 errors larger than the 1280-byte minimum IPv6 MTU are rejected as required by RFC 4443, while informational messages may be fragmented. IsEchoRequest identifies an IPv4 or IPv6 Echo Request, and ReplyEcho constructs its reply directly, preserving identifier, sequence, and data.

IP requests cover valid, reassembled upper-layer protocols that matched no IPConn and are not TCP, UDP, ICMP, or IPv6 No Next Header. A matching raw protocol socket always takes priority. Message exposes the source, destination, protocol number, received hop limit, traffic class, IPv6 flow label, and protocol payload. Reply sends the same protocol number in the reverse direction using the current Config.IP output defaults, while Reject emits IPv4 Protocol Unreachable or IPv6 Parameter Problem. IPv4 protocol number 59 remains an ordinary protocol; No Next Header is special only in IPv6.

UDP, IP, and ICMP handlers run synchronously inside Stack.Write or loopback packet delivery. They may be invoked concurrently by concurrent writers and must return promptly; waiting for traffic that depends on the same delivery call would deadlock it. Request values and payloads returned by request methods are borrowed and valid only during the callback. Reply, ReplyFrom, ReplyIPPacket, and ReplyEcho are repeatable output operations, not terminal actions. A handler may make zero or more reply attempts and then select at most one terminal action. Returning after at least one reply attempt without a terminal action simply completes the request; returning without either applies an implicit Drop. Every request method call must finish before the callback returns.

For UDP, terminal request actions are Accept, Listen, Detach, DetachForReplies, Drop, and Reject; for IP and ICMP they are Detach, DetachForReplies, Drop, and Reject. Both detach methods remove the request from the forwarder's pending set and transfer asynchronous reply ownership to a caller-owned responder. Detach also copies the input snapshot and retains the quote needed by Reject. DetachForReplies avoids those copies when the caller already knows it needs only replies. The ownership and reference directions are:

Stack -> registered Forwarder -> pending callback-scoped Request
                                   |
                                   | Detach or DetachForReplies
                                   v
Caller -> detached Responder -> originating Forwarder state -> Stack

The lower chain is one-way: the responder retains access to its originating forwarder's state and stack for output, diagnostics, and Done, but neither the forwarder nor the stack retains the responder. It is therefore independent of the callback and request lifetime, not independent of forwarder state. The caller may retain it, hand it to another goroutine, or discard it; concurrent access to its mutable snapshot remains the caller's responsibility. A request reply that began before a terminal request action may finish, but a reply begun after it reports ErrForwarderRequestCompleted. Repeated or invalidated terminal request actions report the same error.

A detached UDP, IP, or ICMP responder permits repeated or concurrent reply operations while it is active or restricted to replies. An argument, forwarder, configuration, route, PMTU, or stack error fails only that call and may be followed by another reply or, while active, a terminal action. Local output queue pressure follows the best-effort policy described below. Concurrent calls have no ordering guarantee.

RestrictToReplies irreversibly converts a responder returned by Detach to the same capability set as one returned by DetachForReplies. It releases the responder's references to copied input storage and does not count as Drop or Reject. UDP retains Flow, Reply, ReplyFrom, and Done; IP retains Message metadata, Reply, and Done; ICMP retains Message metadata, Reply, ReplyIPPacket, and Done. Snapshot accessors then return nil payload or packet storage, while ReplyEcho, Drop, and Reject report net.ErrClosed where applicable. Repeated calls, including calls on a responder returned by DetachForReplies, are idempotent no-ops; only a terminal responder causes RestrictToReplies to report net.ErrClosed. The caller must invoke it while no other responder method is running. Reply operations may again run concurrently after it returns. A slice obtained before restriction remains valid and keeps its backing storage live while the caller retains it.

Drop and Reject remain available after any number of responder replies, and exactly one of them may terminate an active responder. They are unavailable after restriction. A reply that began before an active responder's terminal action may finish, while a later reply or terminal action reports net.ErrClosed. The terminal action does not wait for calls already in progress. Neither action is required for resource release because no stack-side collection, timer, or detached-request capacity owns the responder. Discarding it after its final reply releases the caller's reference and contributes no terminal-action diagnostic count. The caller controls its retention, concurrency bound, cancellation, and timeout. Every output call revalidates forwarder closure, the current destination policy, and the return route. Closing the originating forwarder closes the responder's Done channel and makes later output fail with net.ErrClosed; it does not reclaim or mutate any caller-owned snapshot. Configuration changes remain dynamic and are reported by individual calls.

Request-scoped Reply and every Reject action are nonblocking with respect to the outbound packet queue. They use best-effort link output: queue pressure may discard a packet or any queued member of a source-fragmented sequence sent to the external link without turning the action into an error. TCP resets and ICMP errors generated automatically for unhandled local traffic follow the same policy. Ordinary UDP and IP sockets use the bounded-admission policy described below; a stopped device reader never makes their writes wait.

ForwarderInfo.Pending excludes caller-owned responders and handlers that continue running after selecting an action. Accepted counts created TCP/UDP endpoints, Replies counts completed best-effort reply calls, and ReplyErrors counts failed output attempts, including argument and packet validation failures. Calls rejected because a terminal action already completed the request or responder are lifecycle misuse rather than output attempts and do not increment ReplyErrors. Reply counters may exceed Requests when one request produces several replies, and a request may contribute to both reply counters and one terminal-action counter.

Only a forwarded endpoint, request-scoped reply, or detached responder may use an intercepted destination as an output source. Removing that destination's admission by disabling Promiscuous closes affected forwarded connections, removes pending fragments, invalidates callback-scoped requests, and makes later responder output fail current-policy validation. Closing a forwarder stops new fallback requests and invalidates undecided callback-scoped requests without waiting for callbacks to return. An output action already claimed by a request or responder may finish. Accepted TCP and UDP endpoints remain usable until closed or invalidated by configuration.

Active sockets allocate from the IANA dynamic range (49152..65535) first. Only when that range is unavailable for the requested binding or TCP tuple do they fall back to non-privileged ports 1024..49151. Ports below 1024 are never selected automatically but remain available for explicit bindings. TCP combines an RFC 6056-style SipHash offset of the local and remote endpoints with a keyed full-period scan step. Different destinations therefore observe separated sequences, and a collision-free sequence does not revisit a recently closed tuple until it has traversed the complete range. The keys and initial cursors are read from system randomness once when the Stack is created; socket creation does not perform another system-random read. Config.MaxTCPConnections may impose an application-selected resource bound; its zero value does not impose an artificial connection limit. Listener count is controlled only by available memory and explicit application creation.

Config.Routes == nil installs one default route for each configured local address family, or both families for an addressless promiscuous Stack. A non-nil empty route slice deliberately admits only destinations that are themselves local. IPv4-only stacks accept MTUs down to 68; configurations with IPv6 local addresses or output routes require the IPv6 minimum MTU of 1280. Config.AddressProperties supplies the deprecated and temporary state that a netip.Prefix cannot express. Automatic source selection applies the RFC 6724 same-address, scope, deprecation, label, temporary-address, and longest-prefix rules in order. PreferTemporaryAddresses selects temporary IPv6 privacy addresses after label matching; the zero value favors stable addresses. Every property key must identify a configured local address, and Temporary is valid only for IPv6. Config.TCP defines policies inherited by newly created connections and listeners: initial and maximum automatic receive/send buffers, completed and half-open listener queues, congestion control, maximum pacing rate, keepalive, receive-idle timeout, Nagle behavior, TCP user timeout, DSCP bits, and IPv6 Flow Label policy. A zero TCP Flow Label selects a stable keyed label for the connection tuple; a nonzero value fixes the label for new connections. A zero maximum pacing rate is unlimited; nonzero values cap paced data in bytes per second. The initial data burst and control packets are not strictly shaped, so this policy is not a byte-exact traffic shaper. Congestion-control selectors accept registered string names. The built-in names are available as CongestionControlCUBIC, CongestionControlReno, CongestionControlBBR, and CongestionControlBBR3; an empty name selects CUBIC. A programmatic caller may instead supply a local CongestionControlFactory, described below. UpdateConfig applies a changed congestion controller to established connections without an explicit per-connection override. Existing sockets retain the other inherited policies. Receive window scale is selected per connection from the configured receive ceiling, so deliberate small-buffer policies retain window precision while large-BDP connections can use their full automatic maximum. Calling SetReadBuffer or SetWriteBuffer locks that side to the application value and disables its automatic growth, matching the user-locked behavior of operating-system TCP stacks. Automatic growth follows application-consumed and acknowledged bytes per RTT rather than queue size or cwnd alone, so short-RTT scheduler batches do not inflate buffers. SetCongestionControl, SetCongestionControlFactory, SetMaximumPacingRate, and SetTrafficClass provide per-connection overrides. An explicit named or local factory choice is not replaced by later UpdateConfig default changes. Passing zero to SetMaximumPacingRate removes the limit without resetting the controller's path model. BBR pacing groups whole Linux-style send quanta to amortize userspace actor scheduling; a group never exceeds four send quanta, and only one bounded group may be credited ahead of the pacing clock.

Config.UDP and Config.IP define the receive-buffer capacity, path-MTU policy, default TTL/Hop Limit, default TOS/Traffic Class, and IPv6 Flow Label policy inherited by new datagram sockets. Config.IP additionally selects whether new IP protocol sockets read or write complete IP packets. SetReadBuffer, SetPathMTUDiscovery, SetHopLimit, SetTrafficClass, and SetFlowLabel provide per-socket overrides. A zero configured Flow Label uses a stable keyed label for each flow; explicitly setting a socket label to zero disables automatic labeling. Nonzero message control fields override the socket defaults; an explicit zero TOS/Traffic Class or Flow Label encoded in raw OOB data remains distinguishable from an omitted field.

PathMTUDiscovery uses Linux's numeric IP_PMTUDISC_* values. Dont uses confirmed destination PMTU, permits source fragmentation, and leaves IPv4 DF clear. Want also uses destination PMTU, requests DF on a fitting IPv4 packet, and may still fragment locally when needed. Do uses destination PMTU and returns EMSGSIZE instead of fragmenting. Probe ignores destination PMTU, uses the link MTU, requests DF, and returns EMSGSIZE above it. Interface uses the link MTU, leaves DF clear, and rejects an oversized local packet; Omit instead permits source fragmentation. Interface and Omit ignore ICMP PMTU updates for that socket, matching Linux. The zero value is Dont and preserves MIPS's fragmentable datagram default.

Socket operation failures use *net.OpError. errors.Is identifies os.ErrDeadlineExceeded, net.ErrClosed, and syscall errors. Orderly TCP EOF is returned directly as io.EOF, and destination-specific writes on connected UDP or IP sockets retain net.ErrWriteToConnected. Validated asynchronous ICMP errors also match their mapped syscall errors; their details remain available through errors.As to mipstack.ICMPError.

TCPConn.SetLinger provides background graceful close, abortive close, and a bounded wait for acknowledgement. UDPConn.SetReadBuffer changes the receive queue's approximate retained-memory capacity; payload, per-datagram metadata, and queued asynchronous errors share the bound. IPConn applies the same policy. SetReceiveErrors(true) retains reportable asynchronous ICMP errors for nonblocking ReadError; an empty error queue returns EAGAIN. Ordinary reads, UDP writes, and header-included IP writes also report a pending error without removing its ReadError entry; protocol-payload IP writes do not report it. By default, unconnected sockets do not report asynchronous ICMP errors (correlated PMTU updates still apply), while connected sockets report hard errors before queued payloads. Disabling the option clears ReadError entries but preserves an ordinary pending error. ReceiveErrors reports the current mode. UDP and IP writes make one immediate bounded queue-admission attempt. Published backlog is subject to flow-aware replacement. Failure to admit unicast output or an external-link non-unicast copy reports ENOBUFS when enabled and is otherwise a successful message write. The option does not report packets displaced after admission. Receive-side multicast and broadcast loopback copies remain best effort. UDP and IP sockets retain no per-socket transmit queue, so SetWriteBuffer is a validated no-op and a write deadline is checked only before the attempt.

UDP and IP ReadBatch/WriteBatch also accept Linux-compatible message flags. MessageFlagPeek preserves a queued payload but consumes a pending socket error. Like Linux, a successful MessageFlagErrorQueue read consumes its queue entry and may rearm the ordinary error from the next entry. An error-queue read never blocks and returns the quoted failed payload, original destination in Addr, and a Linux sock_extended_err record in OOB. MessageFlagDontWait makes the first batch read nonblocking. Writes are already nonblocking with respect to device capacity, so the flag is accepted without changing their admission result. MessageFlagTruncated requests the complete payload length and, along with MessageFlagControlTruncated, also reports output truncation. Source-fragmented external output admits fragments independently in wire order. Flow-aware overload can discard queued fragments like ordinary link loss, and a later admission failure leaves surviving fragments published. Local loopback output remains all-or-nothing so its reassembler never receives a capacity-truncated datagram.

SocketErrorControlMessage.Parse finds one Linux sock_extended_err record in a possibly compound OOB buffer. MarshalBinary and AppendBinary encode the structured value as one canonical, complete, aligned record; with sufficient destination capacity, AppendBinary can append it to other ancillary data without allocating. The offender address selects IP_RECVERR or IPV6_RECVERR; IPv4-mapped addresses retain their IPv6 sockaddr representation, matching Linux IPv6 socket error queues. Fields not represented by SocketErrorControlMessage, including reserved bytes and sockaddr port, flow-info, and scope fields, are written as zero. An offender must be a valid, unzoned address; Parse rejects Linux AF_UNSPEC offenders because the structured value does not otherwise retain the cmsg address family needed for an unambiguous re-encoding.

An ordinary IPConn exchanges upper-layer protocol payloads. With IPHeaderIncludedOnWrite, IPv4 writes follow Linux IP_HDRINCL: the stack fills a zero source and ID, repairs Total Length and the header checksum, and otherwise preserves the caller's header and payload. Linux IPv6 IPV6_HDRINCL preserves the supplied packet byte-for-byte, and MIPS does the same. The destination argument selects routing independently from the destination stored in either header. IPHeaderIncludedOnRead returns the complete packet after validation and reassembly, retaining IPv4 options and IPv6 extension headers while removing fragmentation state. Caller write buffers and packets queued to different raw sockets have independent ownership.

TCPConn.Info returns a consistent live diagnostic snapshot from the connection actor and retains the final snapshot after close. It includes RFC 9293 state, endpoints, negotiated extensions, RTT/RTO, congestion controller, cwnd and ssthresh, peer/receive windows, bytes in flight, BBR delivery and pacing rates and mode, path MTU and active probe state, buffer occupancy and automatic limits, byte counters, recovery state, inherited keepalive/Nagle/DSCP/Flow Label policies, window-scale values, and connection-local retransmission, PMTU-probe, and spurious-recovery counters. It distinguishes application-limited delivery from host-scheduler-limited delivery and counts material pacing wake delays, which helps distinguish a local runtime stall from a path-bandwidth reduction. The configured maximum pacing rate is reported alongside the effective rate. It also reports the current and peak byte-bounded actor queue occupancy and queue drops, making scheduler or embedding-link backpressure distinguishable from network loss. TCPListener.Info reports current, capacity, and lifetime peak occupancy for the accept and SYN backlogs, along with handshake, SYN-cookie, accept, timeout, and queue-drop counters. UDPConn.Info and IPConn.Info expose endpoint identity, queue occupancy, socket defaults, path MTU for connected sockets, and cumulative receive, receive-drop, and successful socket-write counters. The write counters include default-policy writes silently rejected by local admission and remain cumulative for packets later dropped by link scheduling. Both also report the PMTU-discovery mode, explicit-error mode, queued error count and bytes, and errors dropped by the shared receive-buffer bound. When ICMP error reporting applies, they retain the latest correlated error while open. Closing the socket releases that diagnostic state while preserving cumulative counters. An automatic IPConn Flow Label is reported as zero because raw payload fields may select a different flow on each write; fixed socket labels are reported directly.

UDP message methods use the Linux 64-bit little-endian control-message layout on every host. ReadMsgUDP emits IP_PKTINFO or IPV6_PKTINFO for the packet destination plus TTL/Hop Limit and TOS/Traffic Class, with Linux MSG_TRUNC/MSG_CTRUNC flags. Passing that data to WriteMsgUDP selects the corresponding managed source and output header fields. IPv6 messages also carry Linux IPV6_FLOWINFO. IPConn uses the same ancillary representation.

IPv4ControlMessage and IPv6ControlMessage make that ancillary data structured rather than opaque. Their Parse methods decode OOB returned by a message read; their Marshal methods encode Src, TTL/Hop Limit, and TOS/Traffic Class for a message write. IPv6ControlMessage also exposes the 20-bit Flow Label. Dst is populated while parsing. IfIndex is always zero because MIPS has one embedding link. For IPv4, IPv4ControlMessage.Parse obtains Dst from the IP header destination (ipi_addr), while Marshal uses Src for source selection (ipi_spec_dst).

UDPConn.ReadBatch/WriteBatch and IPConn.ReadBatch/WriteBatch use the same SocketMessage layout as x/net/ipv4 and x/net/ipv6: Buffers contains the scatter/gather payload, OOB contains ancillary data, Addr carries the peer, and N, NN, and Flags receive operation results. A read blocks for its first message and then drains only the currently ready prefix; MessageFlagDontWait makes the first read nonblocking. MessageFlagTruncated and MessageFlagControlTruncated name the fixed Linux result bits returned on every host. Batch operations follow Linux recvmmsg/sendmmsg successful-prefix semantics: an error after one or more completed messages is deferred until the caller retries the unprocessed suffix, whose result fields remain unchanged. Batch writes accept MessageFlagDontWait; other nonzero write flags return EOPNOTSUPP.

Protocol behavior

TCP implements active and passive open, bounded accept and SYN queues, concurrent four-tuple demultiplexing, safe local-port reuse for distinct remote tuples, Linux-style reuse of eligible passive TIME_WAIT tuples using RFC 6191 sequence and timestamp admission checks, and bounded active and TIME_WAIT state. Validated inbound segments wait in a dynamically allocated, byte-bounded FIFO, so idle connections do not pay for a large channel while high-throughput connections are not constrained by an arbitrary segment count. Initial sequence numbers follow RFC 6528: a four-microsecond monotonic counter is added to a SipHash-derived per-four-tuple offset under a 128-bit per-stack secret. Its data path includes bounded send and receive buffers, adaptive RTO with exponential backoff, selectable CUBIC, Reno, BBRv1, and BBRv3 congestion control, window scaling, delayed ACKs, SACK multi-hole recovery with Proportional Rate Reduction, RACK time-based loss detection, tail-loss probes, timestamp negotiation with PAWS, and classic ECN feedback. The receive ACK policy learns the peer's effective segment size without mistaking variable SACK options or application remnants for a smaller MSS. It sends an ACK once unacknowledged data exceeds one learned segment, uses a bounded quick-ACK budget for startup, idle restart, reordering, loss, and ECN feedback, and favors acknowledgement piggybacking for prompt request/reply traffic. Text and FIN carried in a stateful SYN or SYN-ACK are retained through the handshake and processed only after the connection enters ESTABLISHED. SYN-cookie mode remains stateless, so unacknowledged SYN text is accepted only when the peer retransmits it with or after the final ACK.

Loss evidence, SACK/RACK/TLP retransmission selection, PRR inputs, and generic delivery-rate sampling remain owned by TCP rather than an individual congestion controller. Reno, CUBIC, BBRv1, and BBRv3 are separate per-connection implementations behind the public CongestionController event contract; each owns its window policy and private model. Consequently a new controller does not require protocol-specific branches in the TCP actor, and delivery-rate controllers can reuse the common sampler without duplicating TCP sequence or scoreboard logic.

Custom algorithms may be registered process-wide with RegisterCongestionControl and selected by name through TCPSocketDefaults.CongestionControl. Registration is permanent and cannot replace an existing name, so live connections never race an unloaded factory. Programmatic users may instead create an immutable local CongestionControlFactory from a named CongestionControlDefinition with NewCongestionControlFactory, install it in TCPSocketDefaults.CongestionControlFactory, or select it for one live connection with SetCongestionControlFactory. Local factories are not entered in the process registry unless explicitly passed to RegisterCongestionControl, may use the same diagnostic name with different configuration, and compare by pointer identity when a live connection applies an update. The constructor copies and privately retains the definition, so later changes to the caller's value cannot alter the factory. The name and factory fields in TCPSocketDefaults are mutually exclusive.

Every factory invocation receives a by-value CongestionControlContext with the local and remote endpoints and the connection's active, passive, or forwarded role. The context is an immutable identity snapshot rather than a TCPConn, so a factory cannot reenter and deadlock the connection actor. Dynamic MSS, RTT, window, pacing, delivery, loss, and recovery state arrives through events. A factory may capture shared read-only configuration but must create one independent controller per connection; factory calls and controller instances belonging to different connections may run concurrently.

The first controller callback is CongestionEventInitialize; later callbacks are serialized on that connection's actor, reuse the CongestionEvent storage, and must not block or retain it. CongestionEventRelease is the final, observational callback before an implementation is replaced or its connection actor exits. It permits controller-owned cleanup but must return promptly and must not start work that outlives it. Implementations must ignore unknown event types so a newer stack can add observations without changing the one-method controller interface.

CongestionState supplies the Linux-style connection view. Controllers may change cwnd and ssthresh directly during the event types that permit those outputs. They may also select an explicit byte-rate for the common pacer, or declare CongestionControlFeatureCustomPacing when they need to own pacing deadlines and wake accounting. Declaring CongestionControlFeatureDeliveryRate enables per-transmission metadata and a Linux-style sample on ACK events; algorithms that do not need it pay no sampling cost. The sample is a callback-lifetime read-only view exposed through accessors, so sampler internals can evolve without changing controller code. CongestionControlFeatureTransmissionEvents opts into original send and retransmission callbacks and is required by custom pacers. A controller may attach a 64-bit PacketState value to each transmission generation; TCP returns it unchanged through the selected delivery-rate sample. CongestionControlFeatureLossEvents additionally retains that value until the generation is proven lost and reports per-generation SACK, RACK, RTO, and tail-loss-probe recovery events. CongestionRateSample.TailLossProbeACK identifies the Linux-style ambiguous ACK that exactly covers a retransmitted TLP range, allowing a model to preserve delivery signals until later ACK or DSACK evidence resolves the loss. CongestionControlFeatureCustomRecovery opts into PRR, recovery-window, and spurious-undo window decisions. Without it TCP applies its RFC recovery defaults without dispatching those detailed stages; checkpoint and undo notifications remain available to every controller for restoring private state after spurious recovery. CongestionControlFeatureCustomWindowValidation leaves idle and application-limited cwnd validation to a controller that consumes transmission events, as required by model-based algorithms.

Validated network- and host-unreachable feedback for SND.UNA applies RFC 6069 TCP-LD one-step RTO backoff reversion without turning an established connection's soft network error into a hard failure. On a SACK-negotiated connection, only newly reported scoreboard information counts toward RFC 6675 DupAcks; repeated cumulative ACKs and window-probe responses without new SACK data cannot manufacture a loss episode. The receive path preserves three prompt duplicate ACKs and then applies Linux-style bounded SACK compression; the sender gives apparent SACK reneging a short grace period before clearing contradictory scoreboard state and entering timeout recovery. Initial Reno and CUBIC slow start use RFC 9406 HyStart++ and Conservative Slow Start; BBR retains its own Startup model. Eifel timestamps and conservative DSACK accounting detect spurious fast retransmits and timeouts, while RFC 5682 F-RTO detects a spurious timeout without requiring either option. The RFC 4015 response bounds the restored congestion window and makes the RTO more conservative after a spurious timeout is detected. Reno and CUBIC also apply Linux-style RFC 2861 idle and application-limited congestion-window validation; model-based controllers retain their own window policy. TCP also handles overlap-aware receive reassembly, data-bearing zero-window probes, reset validation, deadlines, half-close, FIN states, and TIME_WAIT.

BBR is a byte-scaled implementation of Linux BBRv1. Each original or retransmitted range carries a Linux-style delivery snapshot; ACK processing uses the longer send and acknowledgement phase, a three-candidate ten-round windowed maximum, and a ten-second minimum-RTT filter. Startup, Drain, ProbeBW, and ProbeRTT use the Linux fixed-point gains, randomized ProbeBW phase, ACK aggregation allowance, token-bucket policer detection, idle restart, and first-round packet conservation. SACK and RACK remain responsible for proving loss and selecting retransmissions; transmission-generation accounting keeps merely speculative retransmissions out of BBR's loss model until a replacement generation is independently proven lost, and excludes isolated PLPMTU probe failures. BBR owns pacing and cwnd rather than duplicating the common recovery machinery. Since a Go connection actor does not have kernel fq pacing, a materially late pacing wake is marked locally limited so scheduler delay cannot become a false low path-bandwidth sample. Complete scheduler-limited rounds still participate in Startup plateau detection, preventing sustained host load from trapping the connection in Startup. Pacing retains at most one send quantum of overdue debt, groups at most four whole quanta per actor turn, and credits at most one bounded group ahead of the pacing clock, preventing an unbounded catch-up burst.

BBR3 is an independent byte-scaled implementation of Google's public Linux BBRv3 model. It retains STARTUP, DRAIN, PROBE_BW, and PROBE_RTT; PROBE_BW uses the DOWN, CRUISE, REFILL, and UP phases and the corresponding ACK-feedback state machine. Its path model includes a two-cycle bw_hi maximum, short-term bw_lo and inflight_lo bounds, the robust inflight_hi bound, loss-prefix reconstruction from each transmission generation, loss-based Startup exit, Reno-coexistence and randomized wall-clock probe intervals, ACK aggregation, and the independent five-second ProbeRTT trigger with a ten-second global minimum-RTT window. Recovery undo restores the less restrictive pre-recovery model bounds.

MIPS does not enable BBRv3's low-latency ECN model: classic RFC 3168 ECE does not provide the precise homogeneous CE counts and route-level low-latency ECN eligibility that Linux requires for that model. Classic ECN still uses TCP's ordinary congestion-recovery path. Protective Load Balancing is also absent because an endpoint stack with one embedding link has no kernel route rehash operation. MIPS sends individual TCP segments rather than kernel TSO/GSO skbs, and its connection actor supplies bounded pacing groups in place of Linux fq/EDT. Consequently kernel offload-specific burst sizing is not reproduced, while send-time flight/loss snapshots and all model-visible pacing deadlines remain per connection.

RTT sampling uses packet arrival time rather than actor scheduling time. Its minimum is the Linux-style three-sample running minimum over a 300-second window, so a route change can replace stale path history without retaining an unbounded sample set. When receive work and a protocol timer become ready together, the actor drains the finite receive snapshot that was already queued before servicing the timer. Packets arriving during that turn are excluded, so host scheduling delay cannot manufacture loss and a continuous packet stream cannot starve retransmission, liveness, pacing, or PMTU timers. Transmission timestamps start when packets enter the embedding device queue, and loss timers defer while the original packet still occupies that outbound queue, so link backpressure is not misclassified as network loss.

TCP user timeout follows Linux TCP_USER_TIMEOUT: it applies only in synchronized states, returns ETIMEDOUT, does not change retransmission or keepalive probe timing, and bounds data that remains unacknowledged or unsent behind a zero window. Retransmission does not restart the absolute deadline. When keepalive is enabled, user timeout replaces the probe-count close policy. MIPS treats it as a local socket policy and does not advertise the optional RFC 5482 UTO option.

When the SYN backlog or configured connection capacity is exhausted, passive open uses stateless SYN cookies instead of retaining another half-open connection. A per-stack random key authenticates the complete tuple, client sequence, recent time period, and negotiated options. A valid final ACK reconstructs the conservative MSS and authenticated negotiated options. If that ACK lacks Timestamp, the restored connection disables Timestamp, SACK, window scaling, and ECN. Forged or expired cookies do not allocate a connection.

Validated ICMP Packet Too Big errors maintain a bounded, expiring destination PMTU cache. Error type/code combinations and quoted TCP sequence spans are checked before they can affect transport state. Stack.PathMTU exposes the currently confirmed value, while Stack.ConfirmPathMTU records an application-proven packetization-layer acknowledgement for protocols that manage probing directly. TCP immediately resegments outstanding data and implements RFC 4821 binary-search PLPMTUD when a cached reduction expires. Upward probes carry real data; cumulative ACK confirms success, while only isolated loss proven by SACK suppresses congestion response. Concurrent loss and timeouts remain ordinary congestion and use TCP-friendly probe backoff. MSS changes preserve the required byte/packet congestion units, and successful probes update sibling flows sharing the destination path. A failed path is also reduced by validated ICMP or by IPv4 and IPv6 PMTU black-hole inference after repeated RTOs.

UDP uses the learned MTU for ordinary fragmentation. WritePathMTUProbe and WritePathMTUProbeTo send an explicitly unfragmented packet above the current confirmed PMTU but no larger than the first-hop MTU. Because UDP has no generic acknowledgement, sending alone never raises the PMTU; an application must use its own protected acknowledgement and then call ConfirmPathMTU or ConfirmPathMTUFor, following RFC 8899's packetization-layer contract. Connected UDP sockets select a stable local address and filter inbound remote tuples; unconnected sockets can use arbitrary destinations in one IP family. Both correlate asynchronous ICMP errors with recently used remote endpoints.

IPConn applies the same recent-destination correlation to protocol payload writes. Correlated Packet Too Big updates the shared destination PMTU when the socket's discovery policy permits it. Its explicit probe and confirmation methods follow the same application-acknowledgement contract as UDP.

ICMP echo, unreachable, packet-too-big, IPv4 fragmentation, IPv6 source fragmentation, and bounded IPv4 and IPv6 reassembly are handled internally. Fragment overlap drops the complete datagram. Incomplete sets have count, byte, piece, and lifetime limits. Unsolicited reset, port-unreachable, and echo responses are rate limited, as are RFC 5961 challenge ACKs. ICMPv6 error messages are capped at the IPv6 minimum MTU and active unsupported Routing Headers receive the required Parameter Problem response.

Raw protocol sockets receive reassembled payload copies before the built-in TCP, UDP, or ICMP handler runs. Multiple matching sockets receive independent copies. A listener for an otherwise unknown protocol suppresses Protocol Unreachable while its receive queue accepts or drops matching traffic. ICMPv4Filter follows Linux's 32-bit ICMP_FILTER receive mask, while ICMPv6Filter covers all 256 types defined by RFC 3542. Both may be installed at creation or atomically replaced on an IPConn; packets already queued are not reconsidered. ICMPv6 checksums are verified by default and are inserted for ordinary payload writes. Other raw IPv6 protocols may enable RFC 3542 checksum insertion and verification at an even payload offset through IPv6Checksum. Checksum processing occurs before source fragmentation and after reassembly, and does not alter caller-owned header-included writes.

Stack.Stats returns a lock-free snapshot of active socket counts, separate external-link and loopback packet admissions and queue drops, categorized IP/TCP packet and actor-queue drops, passive handshake, SYN-cookie and accept queue outcomes, retransmission modes, PMTU changes, fragment cleanup, and rate limiting. OutboundQueueDrops covers admission rejection and replacement; LoopbackQueueDrops covers local admission rejection. Neither includes loss after a packet is returned by Stack.Read.

Optional surfaces are arranged for ordinary Go linker reachability rather than build tags. A consumer that only dials TCP and listens for UDP does not retain TCP listener/SYN-cookie code, ReusePort registries, raw IPConn support, message-control helpers, forwarder implementations, or other unreferenced public methods. No package-level registration table or reflection root keeps these APIs alive.

gVisor interoperability

interop/gvisor is an independent Go 1.20 module that links the current checkout directly to the metacubex gVisor fork. Its tests exchange complete L3 packets over an IPv4 MTU matrix from 68 through 9,000 bytes and an IPv6 matrix from 1,280 through 9,000 bytes. Coverage includes bidirectional TCP, UDP, ICMP echo, fragmentation and impairment recovery, every built-in TCP congestion controller, socket errors and ancillary data, transparent Forwarders, broadcast, multicast, PMTU handling, raw IP payload sockets, header-included complete packets, complete-packet reassembly, and Linux-style ICMP error queues. Keeping the tests behind a nested module preserves the root module's standard-library-only dependency graph.

The root go test ./... command does not enter nested modules. Run this suite explicitly with cd interop/gvisor && go test ./... when changing packet or transport behavior.

Scope

MIPS is an endpoint stack, not a general host network stack. IPConn supports protocol-payload and complete-packet raw socket representations, but operating-system file descriptors are deliberately absent. MIPS also does not implement forwarding, NAT, multicast routing, TCP urgent data, or next-hop routing. LocalAddresses controls endpoint ownership and source selection. Routes provides destination admission, longest-prefix selection, metrics, and optional preferred sources, while the lower link remains responsible for gateways, next hops, and L2 neighbor handling. Applications requiring those facilities should use a mature general-purpose userspace stack.

TCP Fast Open is also intentionally absent. A complete server implementation must deliver SYN data before the handshake completes and integrate cookie, SYN-cookie, retransmission, and accept ownership; merely parsing the option or delaying the data until Accept would not implement RFC 7413. The current DialTCP API has no early-data argument, so a partial client implementation would not benefit existing consumers.

License

MIPS is licensed under the Mozilla Public License 2.0. See LICENSE.

Documentation

Overview

Package mipstack implements the mihomo IP stack (MIPS), a small userspace IPv4/IPv6 endpoint stack for applications that exchange complete packets with an L3 link. It implements active and passive TCP, connected and unconnected UDP and IP protocol sockets, and the ICMP behavior required by those transports. Optional forwarders provide transparent handling of otherwise unbound TCP, UDP, ICMP, and other IP protocol traffic.

Index

Constants

View Source
const (
	// CongestionControlCUBIC selects RFC 9438 CUBIC with Reno-friendly growth.
	CongestionControlCUBIC = "cubic"
	// CongestionControlReno selects RFC 5681 Reno congestion avoidance.
	CongestionControlReno = "reno"
	// CongestionControlBBR selects model-based BBR congestion control.
	CongestionControlBBR = "bbr"
	// CongestionControlBBR3 selects Google's loss-bounded BBRv3 model.
	CongestionControlBBR3 = "bbr3"
)
View Source
const (
	// ICMPCodeNone is the zero code used by message types without subcodes.
	ICMPCodeNone = 0

	// ICMPv4TypeEchoReply is the ICMPv4 Echo Reply type.
	ICMPv4TypeEchoReply = 0
	// ICMPv4TypeDestinationUnreachable is the ICMPv4 Destination Unreachable type.
	ICMPv4TypeDestinationUnreachable = 3
	// ICMPv4TypeEchoRequest is the ICMPv4 Echo Request type.
	ICMPv4TypeEchoRequest = 8
	// ICMPv4TypeTimeExceeded is the ICMPv4 Time Exceeded type.
	ICMPv4TypeTimeExceeded = 11
	// ICMPv4TypeParameterProblem is the ICMPv4 Parameter Problem type.
	ICMPv4TypeParameterProblem = 12

	// ICMPv6TypeDestinationUnreachable is the ICMPv6 Destination Unreachable type.
	ICMPv6TypeDestinationUnreachable = 1
	// ICMPv6TypePacketTooBig is the ICMPv6 Packet Too Big type.
	ICMPv6TypePacketTooBig = 2
	// ICMPv6TypeTimeExceeded is the ICMPv6 Time Exceeded type.
	ICMPv6TypeTimeExceeded = 3
	// ICMPv6TypeParameterProblem is the ICMPv6 Parameter Problem type.
	ICMPv6TypeParameterProblem = 4
	// ICMPv6TypeEchoRequest is the ICMPv6 Echo Request type.
	ICMPv6TypeEchoRequest = 128
	// ICMPv6TypeEchoReply is the ICMPv6 Echo Reply type.
	ICMPv6TypeEchoReply = 129

	// ICMPv4DestinationUnreachableCodeNetwork reports an unreachable destination network.
	ICMPv4DestinationUnreachableCodeNetwork = 0
	// ICMPv4DestinationUnreachableCodeHost reports an unreachable destination host.
	ICMPv4DestinationUnreachableCodeHost = 1
	// ICMPv4DestinationUnreachableCodeProtocol reports an unsupported destination protocol.
	ICMPv4DestinationUnreachableCodeProtocol = 2
	// ICMPv4DestinationUnreachableCodePort reports an unreachable destination port.
	ICMPv4DestinationUnreachableCodePort = 3
	// ICMPv4DestinationUnreachableCodeFragmentationNeeded reports a packet that requires fragmentation with DF set.
	ICMPv4DestinationUnreachableCodeFragmentationNeeded = 4
	// ICMPv4DestinationUnreachableCodeSourceRouteFailed reports a failed source route.
	ICMPv4DestinationUnreachableCodeSourceRouteFailed = 5
	// ICMPv4DestinationUnreachableCodeNetworkUnknown reports an unknown destination network.
	ICMPv4DestinationUnreachableCodeNetworkUnknown = 6
	// ICMPv4DestinationUnreachableCodeHostUnknown reports an unknown destination host.
	ICMPv4DestinationUnreachableCodeHostUnknown = 7
	// ICMPv4DestinationUnreachableCodeSourceHostIsolated reports an isolated source host.
	ICMPv4DestinationUnreachableCodeSourceHostIsolated = 8
	// ICMPv4DestinationUnreachableCodeNetworkAdministrativelyProhibited reports a prohibited destination network.
	ICMPv4DestinationUnreachableCodeNetworkAdministrativelyProhibited = 9
	// ICMPv4DestinationUnreachableCodeHostAdministrativelyProhibited reports a prohibited destination host.
	ICMPv4DestinationUnreachableCodeHostAdministrativelyProhibited = 10
	// ICMPv4DestinationUnreachableCodeNetworkUnreachableForTOS reports a network unreachable for the requested TOS.
	ICMPv4DestinationUnreachableCodeNetworkUnreachableForTOS = 11
	// ICMPv4DestinationUnreachableCodeHostUnreachableForTOS reports a host unreachable for the requested TOS.
	ICMPv4DestinationUnreachableCodeHostUnreachableForTOS = 12
	// ICMPv4DestinationUnreachableCodeCommunicationAdministrativelyProhibited reports prohibited communication.
	ICMPv4DestinationUnreachableCodeCommunicationAdministrativelyProhibited = 13
	// ICMPv4DestinationUnreachableCodeHostPrecedenceViolation reports a host-precedence violation.
	ICMPv4DestinationUnreachableCodeHostPrecedenceViolation = 14
	// ICMPv4DestinationUnreachableCodePrecedenceCutoff reports a precedence cutoff.
	ICMPv4DestinationUnreachableCodePrecedenceCutoff = 15

	// ICMPv4TimeExceededCodeTTLInTransit reports an expired IPv4 TTL in transit.
	ICMPv4TimeExceededCodeTTLInTransit = 0
	// ICMPv4TimeExceededCodeFragmentReassembly reports an expired fragment reassembly.
	ICMPv4TimeExceededCodeFragmentReassembly = 1
	// ICMPv4ParameterProblemCodePointer reports an error at the supplied pointer.
	ICMPv4ParameterProblemCodePointer = 0
	// ICMPv4ParameterProblemCodeMissingOption reports a missing required option.
	ICMPv4ParameterProblemCodeMissingOption = 1
	// ICMPv4ParameterProblemCodeBadLength reports an invalid packet length.
	ICMPv4ParameterProblemCodeBadLength = 2

	// ICMPv6DestinationUnreachableCodeNoRoute reports that no route exists.
	ICMPv6DestinationUnreachableCodeNoRoute = 0
	// ICMPv6DestinationUnreachableCodeAdministrativelyProhibited reports prohibited communication.
	ICMPv6DestinationUnreachableCodeAdministrativelyProhibited = 1
	// ICMPv6DestinationUnreachableCodeBeyondSourceScope reports a destination beyond the source address scope.
	ICMPv6DestinationUnreachableCodeBeyondSourceScope = 2
	// ICMPv6DestinationUnreachableCodeAddress reports an unreachable destination address.
	ICMPv6DestinationUnreachableCodeAddress = 3
	// ICMPv6DestinationUnreachableCodePort reports an unreachable destination port.
	ICMPv6DestinationUnreachableCodePort = 4
	// ICMPv6DestinationUnreachableCodeSourceAddressPolicy reports a source-address policy failure.
	ICMPv6DestinationUnreachableCodeSourceAddressPolicy = 5
	// ICMPv6DestinationUnreachableCodeRejectRoute reports a rejected destination route.
	ICMPv6DestinationUnreachableCodeRejectRoute = 6
	// ICMPv6DestinationUnreachableCodeSourceRoutingHeader reports an error in a Source Routing Header.
	ICMPv6DestinationUnreachableCodeSourceRoutingHeader = 7
	// ICMPv6DestinationUnreachableCodeHeadersTooLong reports that processing could not continue because the IPv6 headers were too long.
	ICMPv6DestinationUnreachableCodeHeadersTooLong = 8
	// ICMPv6DestinationUnreachableCodePRoute reports an error in an RPL P-Route.
	ICMPv6DestinationUnreachableCodePRoute = 9

	// ICMPv6TimeExceededCodeHopLimitInTransit reports an expired IPv6 Hop Limit in transit.
	ICMPv6TimeExceededCodeHopLimitInTransit = 0
	// ICMPv6TimeExceededCodeFragmentReassembly reports an expired fragment reassembly.
	ICMPv6TimeExceededCodeFragmentReassembly = 1
	// ICMPv6ParameterProblemCodeErroneousHeaderField reports an erroneous header field.
	ICMPv6ParameterProblemCodeErroneousHeaderField = 0
	// ICMPv6ParameterProblemCodeUnrecognizedNextHeader reports an unrecognized Next Header value.
	ICMPv6ParameterProblemCodeUnrecognizedNextHeader = 1
	// ICMPv6ParameterProblemCodeUnrecognizedOption reports an unrecognized IPv6 option.
	ICMPv6ParameterProblemCodeUnrecognizedOption = 2
	// ICMPv6ParameterProblemCodeIncompleteFirstFragment reports an incomplete first-fragment header chain.
	ICMPv6ParameterProblemCodeIncompleteFirstFragment = 3
	// ICMPv6ParameterProblemCodeSRUpperLayerHeader reports an SR upper-layer header error.
	ICMPv6ParameterProblemCodeSRUpperLayerHeader = 4
	// ICMPv6ParameterProblemCodeUnrecognizedNextHeaderAtIntermediateNode reports an unrecognized Next Header at an intermediate node.
	ICMPv6ParameterProblemCodeUnrecognizedNextHeaderAtIntermediateNode = 5
	// ICMPv6ParameterProblemCodeExtensionHeaderTooBig reports an extension header that exceeds a processing limit.
	ICMPv6ParameterProblemCodeExtensionHeaderTooBig = 6
	// ICMPv6ParameterProblemCodeExtensionHeaderChainTooLong reports an extension-header chain that exceeds a size limit.
	ICMPv6ParameterProblemCodeExtensionHeaderChainTooLong = 7
	// ICMPv6ParameterProblemCodeTooManyExtensionHeaders reports an extension-header count that exceeds a processing limit.
	ICMPv6ParameterProblemCodeTooManyExtensionHeaders = 8
	// ICMPv6ParameterProblemCodeTooManyOptionsInExtensionHeader reports an option count that exceeds a processing limit.
	ICMPv6ParameterProblemCodeTooManyOptionsInExtensionHeader = 9
	// ICMPv6ParameterProblemCodeOptionTooBig reports an option that exceeds a processing limit.
	ICMPv6ParameterProblemCodeOptionTooBig = 10

	// ICMPExtensionClassExtendedInformation identifies RFC 8883 Extended Information objects.
	ICMPExtensionClassExtendedInformation = 4
	// ICMPExtensionExtendedInformationTypePointer identifies an RFC 8883 Pointer object.
	ICMPExtensionExtendedInformationTypePointer = 1
)
View Source
const (
	// ProtocolICMPv4 is the IPv4 Internet Control Message Protocol number.
	ProtocolICMPv4 = 1
	// ProtocolIGMP is the Internet Group Management Protocol number.
	ProtocolIGMP = 2
	// ProtocolTCP is the Transmission Control Protocol number.
	ProtocolTCP = 6
	// ProtocolUDP is the User Datagram Protocol number.
	ProtocolUDP = 17
	// ProtocolESP is the Encapsulating Security Payload protocol number.
	ProtocolESP = 50
	// ProtocolICMPv6 is the IPv6 Internet Control Message Protocol number.
	ProtocolICMPv6 = 58
	// ProtocolNoNextHeader is the IPv6 No Next Header value.
	ProtocolNoNextHeader = 59

	// IPv4HeaderOptionEnd terminates the IPv4 option list.
	IPv4HeaderOptionEnd = 0
	// IPv4HeaderOptionNOP is the one-byte IPv4 No Operation option.
	IPv4HeaderOptionNOP = 1
	// IPv4HeaderOptionRecordRoute records routers traversed by a datagram.
	IPv4HeaderOptionRecordRoute = 7
	// IPv4HeaderOptionTimestamp records router timestamps and optional addresses.
	IPv4HeaderOptionTimestamp = 68
	// IPv4HeaderOptionLooseSourceRoute carries a loose source route.
	IPv4HeaderOptionLooseSourceRoute = 131
	// IPv4HeaderOptionStrictSourceRoute carries a strict source route.
	IPv4HeaderOptionStrictSourceRoute = 137
	// IPv4HeaderOptionRouterAlert requests examination by transit routers.
	IPv4HeaderOptionRouterAlert = 148

	// IPv6ExtensionHeaderHopByHop identifies a Hop-by-Hop Options header.
	IPv6ExtensionHeaderHopByHop = 0
	// IPv6ExtensionHeaderRouting identifies a Routing header.
	IPv6ExtensionHeaderRouting = 43
	// IPv6ExtensionHeaderFragment identifies a Fragment header.
	IPv6ExtensionHeaderFragment = 44
	// IPv6ExtensionHeaderAuthentication identifies an Authentication header.
	IPv6ExtensionHeaderAuthentication = 51
	// IPv6ExtensionHeaderDestination identifies a Destination Options header.
	IPv6ExtensionHeaderDestination = 60
	// IPv6ExtensionHeaderMobility identifies a Mobility header.
	IPv6ExtensionHeaderMobility = 135

	// IPv6ExtensionOptionPad1 is the one-byte IPv6 padding option.
	IPv6ExtensionOptionPad1 = 0
	// IPv6ExtensionOptionPadN is variable-length IPv6 padding.
	IPv6ExtensionOptionPadN = 1
	// IPv6ExtensionOptionRouterAlert requests examination by transit routers.
	IPv6ExtensionOptionRouterAlert = 5
	// IPv6ExtensionOptionJumboPayload carries an IPv6 jumbogram length.
	IPv6ExtensionOptionJumboPayload = 194
	// IPv6ExtensionOptionHomeAddress carries a Mobile IPv6 home address.
	IPv6ExtensionOptionHomeAddress = 201
)
View Source
const (
	// MessageFlagPeek is Linux MSG_PEEK. Ordinary ReadBatch calls copy the oldest
	// queued payload without consuming it. A pending socket error and a
	// MessageFlagErrorQueue result are still consumed, matching Linux recvmsg.
	MessageFlagPeek = 0x02
	// MessageFlagControlTruncated is Linux MSG_CTRUNC. ReadMsg and ReadBatch include
	// it in the result flags when the supplied OOB buffer was too small.
	MessageFlagControlTruncated = 0x08
	// MessageFlagTruncated is Linux MSG_TRUNC. ReadMsg and ReadBatch include it in
	// the result flags when the supplied payload buffers were too small.
	MessageFlagTruncated = 0x20
	// MessageFlagDontWait is Linux MSG_DONTWAIT. Reads return EAGAIN instead of
	// waiting. Datagram writes accept it for compatibility but already use
	// immediate device-queue admission.
	MessageFlagDontWait = 0x40
	// MessageFlagErrorQueue is Linux MSG_ERRQUEUE. ReadBatch reads asynchronous
	// network errors instead of ordinary payloads and never blocks.
	MessageFlagErrorQueue = 0x2000
)
View Source
const (
	// TCPFlagFIN closes one stream direction.
	TCPFlagFIN = 1 << iota
	// TCPFlagSYN synchronizes initial sequence numbers.
	TCPFlagSYN
	// TCPFlagRST resets a connection.
	TCPFlagRST
	// TCPFlagPSH requests prompt delivery to the peer application.
	TCPFlagPSH
	// TCPFlagACK marks AcknowledgmentNumber as valid.
	TCPFlagACK
	// TCPFlagURG marks UrgentPointer as valid.
	TCPFlagURG
	// TCPFlagECE echoes congestion or negotiates ECN on an initial SYN.
	TCPFlagECE
	// TCPFlagCWR acknowledges an ECN congestion response or negotiates ECN.
	TCPFlagCWR
	// TCPFlagNS is the historic ECN Nonce Sum bit, now reserved by RFC 9293.
	TCPFlagNS

	// TCPHeaderOptionEnd terminates the TCP option list.
	TCPHeaderOptionEnd = 0
	// TCPHeaderOptionNOP is the one-byte No-Operation TCP option.
	TCPHeaderOptionNOP = 1
	// TCPHeaderOptionMSS carries a two-byte Maximum Segment Size.
	TCPHeaderOptionMSS = 2
	// TCPHeaderOptionWindowScale carries an unmodified one-byte window scale.
	TCPHeaderOptionWindowScale = 3
	// TCPHeaderOptionSACKPermitted negotiates selective acknowledgments.
	TCPHeaderOptionSACKPermitted = 4
	// TCPHeaderOptionSACK carries one to four selective-acknowledgment blocks.
	TCPHeaderOptionSACK = 5
	// TCPHeaderOptionTimestamp carries TSval and TSecr values.
	TCPHeaderOptionTimestamp = 8
)

Variables

View Source
var (
	// ErrClosed is returned after the stack has been closed.
	ErrClosed = net.ErrClosed
	// ErrNotStarted is returned when packet or socket I/O is attempted before
	// Start.
	ErrNotStarted = errors.New("mipstack: stack is not started")
	// ErrNoPorts reports exhaustion of all automatically allocated,
	// non-privileged ports.
	ErrNoPorts = errors.New("mipstack: no automatic ports available")
	// ErrResourceLimit reports exhaustion of a bounded in-memory socket or
	// protocol resource.
	ErrResourceLimit = errors.New("mipstack: resource limit reached")
	// ErrForwarderRequestCompleted reports an action on a forwarder request that
	// was already accepted, detached, dropped, rejected, invalidated, or whose
	// callback lifetime has ended.
	ErrForwarderRequestCompleted = errors.New("mipstack: forwarder request is already completed")
)

Functions

func AvailableCongestionControls

func AvailableCongestionControls() []string

AvailableCongestionControls returns the registered algorithm names in lexical order. The returned slice is independent of the registry.

func IPTransportChecksum

func IPTransportChecksum(source, destination netip.Addr, protocol int, payload []byte) (uint16, error)

IPTransportChecksum computes an IPv4 or IPv6 pseudo-header checksum over payload. Protocol must be between 0 and 255, payload must not exceed 65535 bytes, and both addresses must be valid, unzoned members of the same family. IPv4-mapped IPv6 addresses are normalized to IPv4. A caller generating a transport header must clear its checksum field first and apply any protocol-specific wire rule; in particular, RFC 768 represents a computed UDP checksum of zero as 0xffff. Computing over a complete valid checksummed payload returns zero.

func IPTransportChecksumParts

func IPTransportChecksumParts(source, destination netip.Addr, protocol int, parts ...[]byte) (uint16, error)

IPTransportChecksumParts computes an IPv4 or IPv6 pseudo-header checksum over parts as though their bytes were concatenated in order. It has the same address, protocol, and protocol-specific wire semantics as IPTransportChecksum. The combined payload must not exceed 65535 bytes.

func InternetChecksum

func InternetChecksum(data []byte) uint16

InternetChecksum computes the RFC 1071 one's-complement checksum of data. A caller generating a checksummed header must clear its checksum field first; computing over a complete valid checksummed region returns zero.

func InternetChecksumParts

func InternetChecksumParts(parts ...[]byte) uint16

InternetChecksumParts computes the RFC 1071 one's-complement checksum of parts as though their bytes were concatenated in order. Empty parts do not disturb word alignment, and an odd trailing byte is paired with the first byte of the next non-empty part. Zero parts computes the checksum of an empty region.

func RegisterCongestionControl

func RegisterCongestionControl(factory *CongestionControlFactory) error

RegisterCongestionControl makes factory available by name to future stack configurations and connections. Registration is concurrency-safe and permanent for the process lifetime, so active connections can keep using the factory without an unregister race. A name cannot be replaced. Factory remains valid for direct local use if registration fails.

Types

type AddressProperties

type AddressProperties struct {
	// Deprecated marks an address past its preferred lifetime. Automatic source
	// selection avoids it when an otherwise usable preferred address exists.
	Deprecated bool
	// Temporary marks an IPv6 privacy address. PreferTemporaryAddresses selects
	// whether automatic source selection favors it over a stable address.
	Temporary bool
}

AddressProperties supplies RFC 6724 source-selection state that cannot be represented by a netip.Prefix alone.

type Config

type Config struct {
	// LocalAddresses lists addresses owned by ordinary sockets and available
	// for source selection and loopback delivery. It may be empty only when
	// Promiscuous is enabled; ordinary sockets then have no usable local family.
	LocalAddresses []netip.Prefix
	// AddressProperties optionally marks configured local addresses as
	// deprecated or temporary for RFC 6724 source selection. Every key must
	// identify an address in LocalAddresses.
	AddressProperties map[netip.Addr]AddressProperties
	// PreferTemporaryAddresses applies RFC 6724 rule 7 in favor of temporary
	// IPv6 privacy addresses. The zero value favors stable public addresses,
	// matching Linux unless per-interface privacy preference is enabled.
	PreferTemporaryAddresses bool
	// Promiscuous admits unicast packets addressed to otherwise nonlocal
	// destinations so protocol forwarders can intercept them. Forwarders do not
	// require Promiscuous for unhandled packets addressed to LocalAddresses.
	// Enabling Promiscuous without a matching forwarder only admits and silently
	// drops nonlocal protocol traffic. Ordinary sockets retain LocalAddresses
	// semantics. Only forwarder-created endpoints and forwarder actions may emit
	// from intercepted addresses. LocalAddresses may therefore be empty; ordinary
	// sockets then cannot bind or select a source.
	Promiscuous bool
	// MTU bounds packets emitted by Read. Zero selects 1500. IPv6 local addresses
	// or output routes require at least 1280.
	MTU uint32
	// Routes optionally restrict admitted unicast destinations and provide a
	// preferred source. Nil installs one default route per configured local
	// address family, or both families for an addressless Promiscuous Stack. A
	// non-nil empty slice installs no routes.
	Routes []Route
	// MaxTCPConnections optionally bounds active, handshaking, and TIME_WAIT
	// connections. Zero leaves the number unbounded; per-listener queues and
	// per-connection buffers remain independently bounded.
	MaxTCPConnections int
	// TCP supplies default socket and listener policies.
	TCP TCPSocketDefaults
	// UDP supplies defaults inherited by new UDP sockets.
	UDP UDPSocketDefaults
	// IP supplies defaults inherited by new IP protocol sockets.
	IP IPSocketDefaults
}

Config configures a Stack.

type CongestionControlContext

type CongestionControlContext struct {
	// LocalAddress is the connection's local TCP endpoint.
	LocalAddress netip.AddrPort
	// RemoteAddress is the connection's peer TCP endpoint.
	RemoteAddress netip.AddrPort
	// Passive reports whether the connection was accepted rather than dialed.
	Passive bool
	// Forwarded reports whether a TCP forwarder accepted the connection for an
	// intercepted destination. Every forwarded connection is also passive.
	Forwarded bool
}

CongestionControlContext identifies the TCP connection for which a controller is being created. It is an immutable, by-value snapshot; dynamic transport state is supplied by CongestionEventInitialize and later events.

type CongestionControlDefinition

type CongestionControlDefinition struct {
	// Name is the diagnostic name and, when registered, the process-registry key.
	Name string
	// New creates one independent controller for the supplied connection.
	New func(CongestionControlContext) CongestionController
	// Features requests optional transport work for the controller.
	Features CongestionControlFeatures
	// SendBufferMultiplier requests a send buffer sized as a multiple of cwnd.
	SendBufferMultiplier uint32
}

CongestionControlDefinition describes a congestion-control implementation before it is validated and frozen into a CongestionControlFactory. New may be called concurrently for different connections, must return promptly, and must return a non-nil, independent controller on every call. It must not retain references to mutable connection state; the context is an immutable value. SendBufferMultiplier requests automatic send-buffer growth to this multiple of cwnd; zero retains the ordinary socket auto-tuning policy.

type CongestionControlFactory

type CongestionControlFactory struct {
	// contains filtered or unexported fields
}

CongestionControlFactory is an immutable, reusable per-connection controller factory. It may be shared by stacks, listeners, and connections; every use invokes its definition's New function to create an independent controller. Factory pointer identity determines whether a live connection's selected implementation changed.

func NewCongestionControlFactory

func NewCongestionControlFactory(definition CongestionControlDefinition) (*CongestionControlFactory, error)

NewCongestionControlFactory validates definition and returns a local factory without registering it process-wide. Definition.Name is used for diagnostics and need not be globally unique. Reuse the returned pointer when successive configurations should identify the same implementation.

func (*CongestionControlFactory) Name

func (f *CongestionControlFactory) Name() string

Name returns the diagnostic name reported by TCPConnInfo. A nil factory has an empty name and is not a valid connection policy.

type CongestionControlFeatures

type CongestionControlFeatures uint32

CongestionControlFeatures declares transport work required by a controller. Unknown bits are rejected when the controller is registered.

const (
	// CongestionControlFeatureDeliveryRate asks TCP to retain per-transmission
	// metadata and include a delivery-rate sample in ACK events.
	CongestionControlFeatureDeliveryRate CongestionControlFeatures = 1 << iota
	// CongestionControlFeatureCustomPacing asks TCP to send pacing query, wake,
	// cancellation, and policy-change events instead of using its window pacer.
	CongestionControlFeatureCustomPacing
	// CongestionControlFeatureTransmissionEvents asks TCP to report original
	// transmissions and retransmissions. Custom pacers require this feature so
	// they can advance their clock only after an actual transmission.
	CongestionControlFeatureTransmissionEvents
	// CongestionControlFeatureCustomRecovery asks TCP to expose recovery-window
	// selection, PRR, partial-ACK, duplicate-ACK, exit, and spurious-undo
	// decisions. TCP applies its RFC defaults when this feature is absent.
	CongestionControlFeatureCustomRecovery
	// CongestionControlFeatureLossEvents asks TCP to retain the opaque packet
	// state returned by transmission events and report each transmission
	// generation that is proven lost. It also enables tail-loss-probe recovery
	// notifications. Transmission events are required so a controller can seed
	// the state associated with each generation.
	CongestionControlFeatureLossEvents
	// CongestionControlFeatureCustomWindowValidation leaves RFC 2861 idle and
	// under-utilization window validation to the controller. Controllers using
	// it must request transmission events so they can observe the first send
	// after an idle interval.
	CongestionControlFeatureCustomWindowValidation
)

type CongestionController

type CongestionController interface {
	// HandleCongestionEvent applies one serialized transport event.
	HandleCongestionEvent(event *CongestionEvent)
}

CongestionController is one connection's congestion-control policy. HandleCongestionEvent is called serially by the connection actor. It must return promptly, must not retain event or references reachable from it, and must ignore event types it does not recognize. The event storage is reused as soon as HandleCongestionEvent returns. The first call for each controller is CongestionEventInitialize and the final call is CongestionEventRelease. An implementation must not start work that outlives the release callback. Different connections own different controller instances and may invoke them concurrently; package-level state therefore requires synchronization.

A controller owns congestion policy and its private model. TCP continues to own acknowledgement validation, loss detection, retransmission selection, SACK/RACK/TLP, PRR accounting, and delivery-rate measurement.

type CongestionDiagnostics

type CongestionDiagnostics struct {
	// DeliveryRate is the current delivery model in bytes per second.
	DeliveryRate uint64
	// PacingRate is the effective paced-data rate in bytes per second.
	PacingRate uint64
	// State is an algorithm-defined short state name.
	State string
	// ApplicationLimited reports the controller's current local limitation.
	ApplicationLimited bool
	// SchedulerLimited reports a material userspace scheduling limitation.
	SchedulerLimited bool
	// SchedulerLimitedEvents counts material userspace pacing wake delays.
	SchedulerLimitedEvents uint64
}

CongestionDiagnostics is the controller-owned portion of TCPConnInfo.

type CongestionEvent

type CongestionEvent struct {
	// Type identifies the valid event payload and permitted outputs.
	Type CongestionEventType
	// Time is ingress time for ACKs and the observation time for other events.
	Time time.Time
	// State is the connection state shared for the callback lifetime.
	State *CongestionState
	// Acknowledged is newly cumulatively acknowledged data in bytes.
	Acknowledged uint32
	// AcknowledgementNumber is the cumulative TCP ACK sequence number.
	AcknowledgementNumber uint32
	// SampleRTT is the current ACK's RTT sample when unambiguous.
	SampleRTT time.Duration
	// RateSample is non-nil on ACK events when delivery-rate sampling is enabled.
	RateSample *CongestionRateSample
	// PacketBytes is the transmission size for packet events.
	PacketBytes int
	// PacketState is controller-owned per-generation state. It is writable only
	// during packet transmission events and read-only during packet-loss and
	// tail-loss-probe events.
	PacketState uint64
	// OutstandingBytes is sequence-space still outstanding before a packet event.
	OutstandingBytes uint32
	// PreviousMaximumSegmentSize is the sender MSS before an MTU event.
	PreviousMaximumSegmentSize int
	// PreviousPhase is the phase before CongestionEventStateChanged.
	PreviousPhase CongestionPhase
	// MarkApplicationLimited asks TCP to mark the current delivery interval.
	MarkApplicationLimited bool
	// Recovery is valid for CongestionEventRecovery.
	Recovery CongestionRecovery
	// Pacing is valid for CongestionEventPacing.
	Pacing CongestionPacing
	// Diagnostics is the output of CongestionEventDiagnostics.
	Diagnostics CongestionDiagnostics
}

CongestionEvent is the reusable event passed to CongestionController. Fields not associated with Type are unspecified and must not be read. State points at the connection's persistent state; controllers may update fields only when the event and payload documentation permit it, and must not replace or retain the pointer. During recovery, only Flight in the selection stage and State.CongestionWindow in custom window stages are outputs. During pacing, only Delay and MarkSchedulerLimited are outputs. Diagnostics is an output only during diagnostic events. RateSample, when non-nil, is a callback-lifetime read-only view and must not be retained. PacketState is an opaque controller output during original-transmission and retransmission events; TCP returns the value unchanged if that generation is selected for a delivery-rate sample, later reported lost, or repaired by a tail-loss probe. A delivery-rate controller may set MarkApplicationLimited during an ACK event to ask TCP to mark the current flight as locally application limited.

type CongestionEventType

type CongestionEventType uint8

CongestionEventType identifies why a controller is being called. New event types may be added without changing CongestionController; implementations must ignore values they do not recognize.

const (
	// CongestionEventUnknown is never sent by TCP.
	CongestionEventUnknown CongestionEventType = iota
	// CongestionEventInitialize seeds a newly established controller and
	// permits updates to its initial cwnd, ssthresh, and pacing-rate policy.
	CongestionEventInitialize
	// CongestionEventACK reports newly delivered data and permits updates to
	// cwnd, ssthresh, and the common pacing-rate policy.
	CongestionEventACK
	// CongestionEventLoss reports entry into loss-based congestion recovery and
	// permits updates to cwnd, ssthresh, and common pacing policy.
	CongestionEventLoss
	// CongestionEventECN reports a new ECN congestion indication and permits
	// updates to cwnd, ssthresh, and common pacing policy.
	CongestionEventECN
	// CongestionEventTimeout reports the first RTO in a recovery episode. TCP
	// accepts updated ssthresh and common pacing policy, then applies the RFC
	// one-MSS timeout window.
	CongestionEventTimeout
	// CongestionEventPacketSent reports an original data transmission to a
	// controller that requested transmission events. It follows delivery-rate
	// snapshot capture and precedes advancement of the common pacer. A custom
	// controller may update cwnd and common pacing policy, for example when
	// leaving a reduced probe state.
	CongestionEventPacketSent
	// CongestionEventPacketRetransmitted reports a retransmission to a controller
	// that requested transmission events, after delivery-rate snapshot capture.
	// It permits updates to common pacing policy.
	CongestionEventPacketRetransmitted
	// CongestionEventRecovery reports a TCP recovery-window transition and
	// permits the outputs documented by CongestionRecovery plus common pacing
	// policy updates.
	CongestionEventRecovery
	// CongestionEventStateChanged reports a congestion-phase transition. State
	// is observational except for persistent pacing policy.
	CongestionEventStateChanged
	// CongestionEventPacing reports custom-pacer work. State is observational.
	CongestionEventPacing
	// CongestionEventMTUChanged reports a new path maximum segment size. State is
	// observational except for persistent pacing policy.
	CongestionEventMTUChanged
	// CongestionEventDiagnostics asks the controller for current diagnostics
	// without changing transport state.
	CongestionEventDiagnostics
	// CongestionEventPacketLost reports one transmission generation newly
	// proven lost by SACK, RACK, or an RTO. PacketBytes and PacketState describe
	// that generation. State is observational.
	CongestionEventPacketLost
	// CongestionEventTailLossProbeRecovered corresponds to Linux
	// CA_EVENT_TLP_RECOVERY: a retransmitted tail-loss probe repaired genuine
	// tail loss rather than merely recovering a lost ACK. PacketBytes describes
	// the repaired range and PacketState is its pre-probe transmission state.
	// State is observational.
	CongestionEventTailLossProbeRecovered
	// CongestionEventRelease is the final callback before TCP drops or replaces
	// a controller. State is observational. Implementations must release any
	// controller-owned resources and must not retain event or references from it.
	CongestionEventRelease
)

type CongestionPacing

type CongestionPacing struct {
	// Operation identifies the custom-pacer interaction.
	Operation CongestionPacingOperation
	// Bytes is the pending transmission size for CongestionPacingQuery.
	Bytes int
	// TransmittedSegments counts original data segments sent by this controller.
	TransmittedSegments uint64
	// Delay is the query result before another transmission may be attempted.
	Delay time.Duration
	// MarkSchedulerLimited asks TCP to mark the current delivery interval.
	MarkSchedulerLimited bool
}

CongestionPacing carries one custom-pacer request and response. Query handlers place the required delay in Delay. Query and wake handlers may set MarkSchedulerLimited when local scheduling, rather than the path, delayed the current flight.

type CongestionPacingOperation

type CongestionPacingOperation uint8

CongestionPacingOperation identifies a custom-pacing interaction.

const (
	// CongestionPacingUnknown is never sent by TCP.
	CongestionPacingUnknown CongestionPacingOperation = iota
	// CongestionPacingQuery asks how long the next transmission must wait.
	CongestionPacingQuery
	// CongestionPacingWake reports an actor wake requested by the controller.
	CongestionPacingWake
	// CongestionPacingCancel invalidates a pending pacing wake.
	CongestionPacingCancel
	// CongestionPacingPolicyChanged reports a new maximum pacing rate.
	CongestionPacingPolicyChanged
)

type CongestionPhase

type CongestionPhase uint8

CongestionPhase is TCP's high-level congestion state, corresponding to the state exposed to Linux congestion-control modules.

const (
	// CongestionPhaseUnknown is used before controller initialization.
	CongestionPhaseUnknown CongestionPhase = iota
	// CongestionPhaseOpen has no active congestion response.
	CongestionPhaseOpen
	// CongestionPhaseDisorder has duplicate or selective ACK evidence but no
	// proven loss. TCP may add finer-grained disorder notifications later.
	CongestionPhaseDisorder
	// CongestionPhaseCWR is responding to ECN congestion.
	CongestionPhaseCWR
	// CongestionPhaseRecovery is fast loss recovery.
	CongestionPhaseRecovery
	// CongestionPhaseLoss is retransmission-timeout recovery.
	CongestionPhaseLoss
)

type CongestionRateSample

type CongestionRateSample struct {
	// contains filtered or unexported fields
}

CongestionRateSample is one ACK's read-only delivery-rate observation. Its interval is the longer of the send and acknowledgement phases. Valid reports false when TCP could not form an unambiguous rate sample; ACK and loss accounting remain usable. Rates use bytes because mipstack's SACK scoreboard is byte exact.

TCP owns this value. A controller may read it only while handling the event that supplied it and must not retain its pointer. Accessors deliberately hide transport-only transmission-selection metadata so future sampler changes do not alter the public congestion-control contract.

func (*CongestionRateSample) ACKDelayed

func (s *CongestionRateSample) ACKDelayed() bool

ACKDelayed reports a lone runt sample likely delayed by the receiver.

func (*CongestionRateSample) ACKTime

func (s *CongestionRateSample) ACKTime() time.Time

ACKTime returns packet ingress time rather than later processing time.

func (*CongestionRateSample) AcknowledgedBytes

func (s *CongestionRateSample) AcknowledgedBytes() uint32

AcknowledgedBytes returns the cumulative ACK byte advance.

func (*CongestionRateSample) ApplicationLimited

func (s *CongestionRateSample) ApplicationLimited() bool

ApplicationLimited reports a sender bubble in the sampled interval.

func (*CongestionRateSample) BytesInFlight

func (s *CongestionRateSample) BytesInFlight() uint32

BytesInFlight returns flight after ACK processing and transmissions caused by the same ACK.

func (*CongestionRateSample) DeliveredBytes

func (s *CongestionRateSample) DeliveredBytes() uint32

DeliveredBytes returns bytes delivered since the sampled transmission's delivery snapshot.

func (*CongestionRateSample) InFastRecovery

func (s *CongestionRateSample) InFastRecovery() bool

InFastRecovery reports fast recovery specifically.

func (*CongestionRateSample) InRecovery

func (s *CongestionRateSample) InRecovery() bool

InRecovery reports fast- or timeout-recovery processing for this ACK.

func (*CongestionRateSample) Interval

func (s *CongestionRateSample) Interval() time.Duration

Interval returns the rate interval selected from the send and ACK phases.

func (*CongestionRateSample) LostBytes

func (s *CongestionRateSample) LostBytes() uint64

LostBytes returns the bytes newly proven lost by this sample.

func (*CongestionRateSample) PacketState

func (s *CongestionRateSample) PacketState() uint64

PacketState returns the opaque state produced by the transmission event for the range selected to form this sample. It is zero when the controller did not request transmission events or the range predates a controller change.

func (*CongestionRateSample) PriorBytesInFlight

func (s *CongestionRateSample) PriorBytesInFlight() uint32

PriorBytesInFlight returns flight immediately before ACK processing.

func (*CongestionRateSample) PriorDeliveredBytes

func (s *CongestionRateSample) PriorDeliveredBytes() uint64

PriorDeliveredBytes returns the cumulative delivered-byte count captured when the sampled range was transmitted.

func (*CongestionRateSample) RTT

RTT returns the selected range's round-trip time when unambiguous.

func (*CongestionRateSample) Retransmitted

func (s *CongestionRateSample) Retransmitted() bool

Retransmitted reports ambiguous delivery through a retransmitted range.

func (*CongestionRateSample) SchedulerLimited

func (s *CongestionRateSample) SchedulerLimited() bool

SchedulerLimited reports material local scheduling delay in the interval.

func (*CongestionRateSample) SmoothedRTT

func (s *CongestionRateSample) SmoothedRTT() time.Duration

SmoothedRTT returns the current RFC 6298 smoothed round-trip time.

func (*CongestionRateSample) TailLossProbeACK

func (s *CongestionRateSample) TailLossProbeACK() bool

TailLossProbeACK reports that this ACK exactly covers a retransmitted tail-loss probe whose original and probe deliveries cannot yet be distinguished. Model-based controllers can use it to retain round-local delivery signals, matching Linux rate_sample.is_acking_tlp_retrans_seq.

func (*CongestionRateSample) Valid

func (s *CongestionRateSample) Valid() bool

Valid reports whether DeliveredBytes divided by Interval is a usable rate.

type CongestionRecovery

type CongestionRecovery struct {
	// Stage identifies the current recovery interaction.
	Stage CongestionRecoveryStage
	// SACK reports whether selective acknowledgement recovery is active.
	SACK bool
	// OrdinaryFlight is TCP's ordinary bytes-in-flight estimate.
	OrdinaryFlight uint32
	// LossFlight is the RFC loss-recovery flight estimate.
	LossFlight uint32
	// Flight is the mutable result of CongestionRecoverySelectFlight and the
	// selected or current recovery flight during later stages.
	Flight uint32
	// PreviousWindow is cwnd before TCP installs ProposedWindow.
	PreviousWindow uint32
	// Acknowledged is the byte advance for a partial ACK.
	Acknowledged uint32
	// ProposedWindow is TCP's RFC-default cwnd for this stage.
	ProposedWindow uint32
}

CongestionRecovery describes a transport-owned recovery transition. TCP places its RFC-default result in State.CongestionWindow before dispatch; controllers with CongestionControlFeatureCustomRecovery may replace it. CongestionRecoveryUndo also permits replacing State.SlowStartThreshold. Flight is initialized to the transport default during CongestionRecoverySelectFlight and is the only mutable field of that stage.

type CongestionRecoveryStage

type CongestionRecoveryStage uint8

CongestionRecoveryStage identifies one transport-owned recovery step.

const (
	// CongestionRecoveryUnknown is never sent by TCP.
	CongestionRecoveryUnknown CongestionRecoveryStage = iota
	// CongestionRecoveryCheckpoint precedes a recoverable congestion signal.
	CongestionRecoveryCheckpoint
	// CongestionRecoverySelectFlight selects the flight used to enter recovery.
	CongestionRecoverySelectFlight
	// CongestionRecoveryEnter applies the initial fast-recovery window.
	CongestionRecoveryEnter
	// CongestionRecoveryPRR applies a PRR-proposed window.
	CongestionRecoveryPRR
	// CongestionRecoveryExit leaves fast recovery.
	CongestionRecoveryExit
	// CongestionRecoveryPartialACK handles a NewReno partial ACK.
	CongestionRecoveryPartialACK
	// CongestionRecoveryDuplicateACK applies non-SACK duplicate-ACK inflation.
	CongestionRecoveryDuplicateACK
	// CongestionRecoveryUndo reports that recovery was proven spurious. A
	// custom-recovery controller may replace State.CongestionWindow and
	// State.SlowStartThreshold; TCP otherwise retains its RFC response.
	CongestionRecoveryUndo
)

type CongestionState

type CongestionState struct {
	// CongestionWindow is cwnd in bytes.
	CongestionWindow uint32
	// SlowStartThreshold is ssthresh in bytes.
	SlowStartThreshold uint32
	// BytesInFlight is the transport's current congestion flight in bytes.
	BytesInFlight uint32
	// MaximumSegmentSize is the current sender MSS in bytes.
	MaximumSegmentSize int
	// SmoothedRTT is the current RFC 6298 smoothed round-trip time.
	SmoothedRTT time.Duration
	// MinimumRTT is the current transport minimum-RTT estimate.
	MinimumRTT time.Duration
	// UsePacingRate selects PacingRate instead of TCP's window-derived rate.
	UsePacingRate bool
	// PacingRate is the requested or current model rate in bytes per second.
	PacingRate uint64
	// MaximumPacingRate is the socket pacing ceiling in bytes per second.
	MaximumPacingRate uint64
	// DeliveredBytes is the cumulative delivery-rate sampler byte count.
	DeliveredBytes uint64
	// LostBytes is the cumulative transport-proven lost byte count.
	LostBytes uint64
	// ApplicationLimited reports an active application-limited interval.
	ApplicationLimited bool
	// SchedulerLimited reports an active local-scheduler-limited interval.
	SchedulerLimited bool
	// SchedulerLimitedEvents counts material local pacing wake delays.
	SchedulerLimitedEvents uint64
	// Phase is TCP's current congestion-control phase.
	Phase CongestionPhase
}

CongestionState is the transport snapshot visible to a controller. TCP applies changes to CongestionWindow, SlowStartThreshold, UsePacingRate, and PacingRate only for event types that explicitly permit them; the remaining fields are read-only. Pacing policy persists between callbacks. When UsePacingRate is false, TCP derives pacing from cwnd and SRTT. Setting it true makes the common pacer use PacingRate bytes per second. A zero duration or counter means that the value is not yet known.

type DatagramSocketDefaults

type DatagramSocketDefaults struct {
	// ReceiveBuffer is the approximate retained-memory receive capacity.
	ReceiveBuffer int
	// ReceiveErrors retains reportable asynchronous network errors for ReadError.
	// When disabled, unconnected sockets do not report them; connected sockets
	// report hard errors before queued payloads on ordinary reads. UDP and
	// header-included IP writes may also return a pending error. IP writes of
	// protocol payloads leave it for reads. When enabled, eligible soft errors
	// follow the same rules, and ordinary operations do not remove their
	// ReadError entries. It also makes failed immediate admission of unicast or
	// external-link multicast/broadcast output return ENOBUFS.
	// Later packet displacement is not reported. Receive-side non-unicast
	// loopback copies remain best effort.
	ReceiveErrors bool
	// PathMTUDiscovery selects the Linux-compatible source-fragmentation and
	// destination-PMTU policy. The zero value is PathMTUDiscoveryDont.
	PathMTUDiscovery PathMTUDiscovery
	// HopLimit is the default IPv4 TTL or IPv6 Hop Limit. Zero selects 64.
	HopLimit int
	// MulticastHopLimit is the default IPv4 multicast TTL or IPv6 multicast
	// Hop Limit. Zero selects the socket-compatible default of one hop.
	MulticastHopLimit int
	// DisableMulticastLoopback disables delivery of transmitted multicast
	// packets to matching local memberships. Loopback is enabled by default,
	// matching IP_MULTICAST_LOOP and IPV6_MULTICAST_LOOP.
	DisableMulticastLoopback bool
	// DisableBroadcast clears the SO_BROADCAST-equivalent permission inherited
	// by new sockets. Broadcast is enabled by default, matching Go's net UDP
	// sockets on supported operating systems.
	DisableBroadcast bool
	// TrafficClass is the default IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass uint8
	// FlowLabel is the default IPv6 Flow Label. Zero selects a stable automatic
	// label for each destination flow.
	FlowLabel uint32
}

DatagramSocketDefaults configures policies shared by newly created UDP and IP protocol sockets. Zero fields retain the package defaults.

type Dialer

type Dialer struct {
	// Options contains creation policies applied before connecting. The call
	// reads but does not retain the slice. Repeated option kinds use the last
	// value.
	Options []SocketOption
}

Dialer contains options applied before a socket is connected. The zero value is ready for use.

func (*Dialer) DialIP

func (dialer *Dialer) DialIP(ctx context.Context, stack *Stack, network string, source, remote netip.Addr) (net.Conn, error)

DialIP creates a connected IP protocol socket through stack.

func (*Dialer) DialTCP

func (dialer *Dialer) DialTCP(ctx context.Context, stack *Stack, network string, source, remote netip.AddrPort) (net.Conn, error)

DialTCP establishes a TCP connection through stack.

func (*Dialer) DialUDP

func (dialer *Dialer) DialUDP(ctx context.Context, stack *Stack, network string, source, remote netip.AddrPort) (net.Conn, error)

DialUDP creates a connected UDP socket through stack.

type ForwarderFlow

type ForwarderFlow struct {
	// Source is the remote endpoint that sent the packet.
	Source netip.AddrPort
	// Destination is the original local endpoint from the packet.
	Destination netip.AddrPort
}

ForwarderFlow identifies one inbound TCP or UDP four-tuple. Source is the remote endpoint and Destination is the original packet destination.

type ForwarderInfo

type ForwarderInfo struct {
	// Closed reports whether this forwarder has been unregistered and closed.
	Closed bool
	// Pending is the number of callback-scoped requests that have not completed
	// or transferred traffic ownership. It excludes detached responders and
	// handlers that continue running after selecting an action.
	Pending int
	// MaxInFlight is the TCP pending-request bound, or zero for protocols that
	// do not retain asynchronous requests.
	MaxInFlight int
	// Requests counts requests delivered to the handler.
	Requests uint64
	// Accepted counts successfully created TCP connections, connected UDP
	// flows, and unconnected UDP listeners.
	Accepted uint64
	// Replies counts completed best-effort request and responder reply calls.
	Replies uint64
	// ReplyErrors counts argument, state, and output failures after a Reply call
	// has entered its request or responder lifetime. Calls rejected because a
	// terminal action already completed that lifetime are not output attempts.
	ReplyErrors uint64
	// Dropped counts explicit, implicit, invalidated, and pending-request
	// capacity drops.
	Dropped uint64
	// Rejected counts explicit protocol rejection decisions.
	Rejected uint64
}

ForwarderInfo is a diagnostic snapshot of endpoint, request, and reply activity for one protocol forwarder.

type ICMPError

type ICMPError struct {
	// Reporter is the router or destination that generated the error.
	Reporter netip.Addr
	// Type is the wire ICMP error type.
	Type byte
	// Code retains the wire subtype within Type.
	Code byte
	// MTU is present for fragmentation-needed and packet-too-big errors. It is
	// also consumed by ICMPError.ICMPMessage when the error type defines that
	// field.
	MTU uint32
	// Pointer is present for IPv4 or IPv6 Parameter Problem errors. It is also
	// consumed by ICMPError.ICMPMessage when the error type defines that field.
	Pointer uint32
	// Extensions contains the encoded RFC 4884 extension-object sequence,
	// excluding the four-byte Extension Header. Parsed errors borrow the ICMP
	// message body. ICMPError.ICMPMessage validates and copies this storage.
	Extensions []byte
	// QuotedSource is the source of the failed original packet. It is derived
	// from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedSource netip.Addr
	// QuotedTarget is the destination of the failed original packet. It is
	// derived from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedTarget netip.Addr
	// QuotedProtocol is the final protocol identified in the available quote.
	// It is derived from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedProtocol byte
	// QuotedPacket contains the available original IP packet, including its IP
	// header. ICMPError.ICMPMessage copies it into the constructed error.
	// QuotedPayload aliases its upper-layer suffix when both are present.
	QuotedPacket []byte
	// QuotedPayload contains the available original transport header bytes. It
	// is derived from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedPayload []byte
	// QuotedSourcePort is the original TCP or UDP source port when present. It is
	// derived from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedSourcePort uint16
	// QuotedTargetPort is the original TCP or UDP destination port when present.
	// It is derived from QuotedPacket and ignored by ICMPError.ICMPMessage.
	QuotedTargetPort uint16
}

ICMPError describes a validated remote network error.

func (ICMPError) Error

func (e ICMPError) Error() string

Error formats the remote ICMP failure.

func (ICMPError) ExtensionObjects

func (e ICMPError) ExtensionObjects() ([]ICMPExtensionObject, error)

ExtensionObjects parses the RFC 4884 object sequence. The returned slice owns its descriptors, but every Data field borrows e.Extensions. An error reports malformed object framing. An ICMPError with no extensions returns a nil slice and nil error.

func (ICMPError) ICMPMessage

func (e ICMPError) ICMPMessage(destination netip.Addr) (ICMPMessage, error)

ICMPMessage constructs the semantic ICMP error represented by e for destination. Reporter and destination select the address family; Type, Code, MTU, Pointer, QuotedPacket, and Extensions provide the wire fields. It validates the possibly truncated quote and extension objects, copies both byte slices, and ignores the fields derived from the quote. Routing, source ownership, rate limits, quote truncation, and recursive-error suppression remain Stack transmission policy.

func (*ICMPError) SetExtensionObjects

func (e *ICMPError) SetExtensionObjects(objects []ICMPExtensionObject) error

SetExtensionObjects replaces Extensions with an encoded RFC 4884 object sequence. Each Data length must be a multiple of four bytes. The method preserves unknown classes and types, duplicates, and order, copies every input Data slice, and leaves e unchanged on failure. An empty sequence clears Extensions; an encoded Extension Structure is generated only when at least one object is present.

func (ICMPError) Unwrap

func (e ICMPError) Unwrap() error

Unwrap returns the syscall error associated with a valid ICMP type and code. It uses the nearest socket errno where the platform has no exact equivalent. The Linux errno carried by MSG_ERRQUEUE is a separate, fixed wire format.

type ICMPExtensionObject

type ICMPExtensionObject struct {
	// Class is the object's Class-Num.
	Class uint8
	// Type is the class-specific C-Type.
	Type uint8
	// Data is the object payload following Length, Class-Num, and C-Type.
	Data []byte
}

ICMPExtensionObject is one RFC 4884 extension object in semantic wire order. Data excludes the four-byte object header. ExtensionObjects returns Data slices that borrow ICMPError.Extensions, while SetExtensionObjects copies every Data slice. Unknown classes and types, duplicates, and order are preserved.

func (ICMPExtensionObject) Pointer

func (o ICMPExtensionObject) Pointer() (uint32, bool)

Pointer returns the byte offset carried by an RFC 8883 Extended Information Pointer object.

func (*ICMPExtensionObject) SetPointer

func (o *ICMPExtensionObject) SetPointer(pointer uint32)

SetPointer replaces o with an RFC 8883 Extended Information Pointer object.

type ICMPForwarder

type ICMPForwarder struct {
	// contains filtered or unexported fields
}

ICMPForwarder owns the single fallback ICMP handler installed on a stack. It may be installed before or after Stack.Start.

func NewICMPForwarder

func NewICMPForwarder(stack *Stack, options ICMPForwarderOptions, handler ICMPForwarderHandler) (*ICMPForwarder, error)

NewICMPForwarder installs a fallback handler for otherwise unhandled ICMP messages. Only one ICMP forwarder may be active per stack. Promiscuous mode is not required for unhandled traffic addressed to LocalAddresses; Config.Promiscuous is required only for nonlocal destination addresses. Installing a forwarder does not start the stack.

func (*ICMPForwarder) Close

func (f *ICMPForwarder) Close() error

Close removes the ICMP fallback handler and invalidates undecided requests. It does not wait for running handlers to return. An action already claimed by a handler or detached responder may finish.

func (ICMPForwarder) Done

func (f ICMPForwarder) Done() <-chan struct{}

Done is closed when the forwarder is closed directly or by Stack.Close.

func (*ICMPForwarder) Info

func (f *ICMPForwarder) Info() ForwarderInfo

Info returns one ICMP forwarder diagnostic snapshot.

type ICMPForwarderHandler

type ICMPForwarderHandler func(*ICMPForwarderRequest)

ICMPForwarderHandler processes one checksum-valid ICMP message not consumed by the stack's built-in echo or asynchronous-error handling. Its synchronous, concurrent-call, and blocking rules are the same as UDPForwarderHandler. The handler must call Detach, DetachForReplies, Drop, Reject, or at least one Reply before returning; returning without an action drops the message. Reply may be repeated and does not prevent a later terminal action. Except for Detach and DetachForReplies, every action and Reply call must finish before the handler returns. The request and Message.Payload must not be retained after that point, but a responder returned by either detach method may outlive the callback. Responder output remains subject to the originating forwarder's state.

type ICMPForwarderMessage

type ICMPForwarderMessage struct {
	// Source is the sender of the ICMP message.
	Source netip.Addr
	// Destination is the original packet destination.
	Destination netip.Addr
	// Type and Code retain the wire classification. Unknown or unassigned
	// values can reach an ICMP forwarder when no built-in handler consumes them.
	Type uint8
	// Code retains the wire subtype within Type.
	Code uint8
	// Payload contains the complete ICMP header and body.
	Payload []byte
}

ICMPForwarderMessage describes one checksum-validated, reassembled ICMP protocol message. Payload contains the complete ICMP header and body. Its ownership and lifetime are specified by the method that returned the message.

func (ICMPForwarderMessage) ICMPMessage

func (m ICMPForwarderMessage) ICMPMessage() (ICMPMessage, error)

ICMPMessage validates and decodes the current wire message. Body aliases Payload[4:] and inherits Payload's ownership and lifetime. It reports syscall.EINVAL when the message is incomplete, its checksum is invalid, or Type and Code no longer agree with Payload.

func (ICMPForwarderMessage) IsEchoRequest

func (m ICMPForwarderMessage) IsEchoRequest() bool

IsEchoRequest reports whether the message is a complete IPv4 or IPv6 Echo Request whose Type and Code fields agree with Payload. Source and Destination must identify the same address family; IPv4-mapped addresses select IPv4.

func (*ICMPForwarderMessage) SetICMPMessage

func (m *ICMPForwarderMessage) SetICMPMessage(message ICMPMessage) error

SetICMPMessage replaces m with the complete wire encoding of message. It validates message before changing m, normalizes IPv4-mapped addresses, and reuses the current Payload capacity when possible. Message.Body may alias Payload; a successful call does not retain any other input storage. On failure m is unchanged.

type ICMPForwarderOptions

type ICMPForwarderOptions struct{}

ICMPForwarderOptions reserves ICMP interception policy for future extension.

type ICMPForwarderRequest

type ICMPForwarderRequest struct {
	// contains filtered or unexported fields
}

ICMPForwarderRequest is one checksum-valid ICMP message not consumed by built-in protocol handling. The handler may reply repeatedly before selecting at most one terminal action.

func (*ICMPForwarderRequest) Detach

Detach transfers one ICMP request out of the synchronous handler lifetime. On success it removes the request from the forwarder's pending set and returns a caller-owned message snapshot. The responder points back to the originating forwarder's state for output and Done, but the forwarder does not retain the responder or impose a capacity or timeout. The caller may hand it to another goroutine or discard it without a terminal action. Detach itself is the request's action and consumes the request even when it returns an error. It may be called after any number of Reply, ReplyIPPacket, or ReplyEcho attempts; the responder remains available for further replies.

func (*ICMPForwarderRequest) DetachForReplies

func (r *ICMPForwarderRequest) DetachForReplies() (*ICMPForwarderResponder, error)

DetachForReplies transfers the ICMP request into an asynchronous responder that retains only Message metadata, Reply, ReplyIPPacket, and Done. Message().Payload and IPPacket return nil, while ReplyEcho, Reject, and Drop report net.ErrClosed. Unlike Detach, it does not copy the triggering packet or retain an ICMP rejection quote. The caller owns the responder and may discard it without a terminal action. RestrictToReplies is an idempotent no-op on the returned responder.

func (*ICMPForwarderRequest) Drop

func (r *ICMPForwarderRequest) Drop() error

Drop consumes the ICMP message without packet I/O. It may follow any number of Reply, ReplyIPPacket, or ReplyEcho attempts, may wait briefly for forwarder bookkeeping, and does not wait for the network or output queue.

func (*ICMPForwarderRequest) IPPacket

func (r *ICMPForwarderRequest) IPPacket() []byte

IPPacket returns the complete, reassembled L3 packet that contains Message. The returned slice aliases packet-delivery storage, is read-only, and is valid only until the handler returns. Message().Payload aliases the ICMP portion of the same packet.

func (*ICMPForwarderRequest) Message

Message returns the checksum-validated ICMP message presented to the handler. Payload aliases packet-delivery storage, must not be modified, and is valid only until the handler returns. Message does not select an action.

func (*ICMPForwarderRequest) Reject

func (r *ICMPForwarderRequest) Reject() error

Reject consumes the ICMP message and emits an administratively prohibited response when ICMP rules permit an error response. It does not wait for outbound capacity, and local output congestion may discard the response without error. It reports syscall.EADDRNOTAVAIL when the intercepted destination is no longer admitted and syscall.ENETUNREACH when no return route remains. The rejection decision remains terminal and may follow any number of reply attempts.

func (*ICMPForwarderRequest) Reply

func (r *ICMPForwarderRequest) Reply(payload []byte) error

Reply sends a complete ICMP protocol message from Destination to Source. The method may be called repeatedly or concurrently, including before a later terminal action, but every call must finish before the handler returns. The stack copies payload, recalculates its checksum, and makes one immediate best-effort output attempt. Local output congestion may discard the message or any of its source fragments without error. Other errors may be retried.

func (*ICMPForwarderRequest) ReplyEcho

func (r *ICMPForwarderRequest) ReplyEcho() error

ReplyEcho copies the triggering Echo Request into an Echo Reply, preserving its identifier, sequence, and data. It reports syscall.EINVAL when the triggering message is not an IPv4 or IPv6 Echo Request. Like Reply, it may be retried and does not prevent a later terminal action; every call must finish before the handler returns.

func (*ICMPForwarderRequest) ReplyIPPacket

func (r *ICMPForwarderRequest) ReplyIPPacket(packet []byte) error

ReplyIPPacket sends a complete IPv4 ICMP or IPv6 ICMP packet whose final destination is the triggering packet's source. It is a restricted header-included ICMP operation, not arbitrary packet injection. Its source may be any valid same-family address; it need not belong to LocalAddresses and is not classified as unicast, multicast, or broadcast here. MIPS copies packet, normalizes its IP length and outer IPv4 and ICMP checksums, preserves other legal header fields and extension headers, and source-fragments it when permitted. A non-atomic input fragment is invalid. An IPv6 atomic Fragment header is preserved when the packet fits; when fragmentation is required, it is replaced by the emitted fragment sequence instead of nesting another header. Output does not wait for capacity; local congestion may discard the packet or any of its source fragments without error. The method may be retried after other validation or output failures and does not prevent a later terminal action.

type ICMPForwarderResponder

type ICMPForwarderResponder struct {
	// contains filtered or unexported fields
}

ICMPForwarderResponder owns one detached message. A responder returned by Detach owns its packet snapshot; DetachForReplies omits that snapshot. It may outlive the handler because the caller, not the forwarder, owns it. It retains access to the originating forwarder's state for output, diagnostics, and Done; neither the forwarder nor the stack retains the responder. Reply and ReplyIPPacket calls may be repeated while active or restricted to replies; ReplyEcho, Reject, and Drop are available only while active. The responder may be discarded without a terminal action.

func (*ICMPForwarderResponder) Done

func (r *ICMPForwarderResponder) Done() <-chan struct{}

Done is closed when the originating protocol forwarder is closed directly or by Stack.Close.

func (*ICMPForwarderResponder) Drop

func (r *ICMPForwarderResponder) Drop() error

Drop terminates the detached input without packet I/O. It remains valid after replies while the responder is active. It reports net.ErrClosed when the responder is already terminal or restricted to replies, including one returned by DetachForReplies.

func (*ICMPForwarderResponder) IPPacket

func (r *ICMPForwarderResponder) IPPacket() []byte

IPPacket returns the complete, reassembled packet snapshot retained by Detach, or nil after DetachForReplies or RestrictToReplies. The caller owns a returned slice and may retain or modify it, but must synchronize concurrent access. Message().Payload aliases its ICMP region.

func (*ICMPForwarderResponder) Message

Message returns the detached ICMP metadata and independently owned payload. Payload is nil after DetachForReplies or RestrictToReplies. The caller may retain or modify a returned payload and must synchronize concurrent access. Type and Code remain the original classification; changing the first two Payload bytes makes the snapshot inconsistent and causes ReplyEcho to report syscall.EINVAL.

func (*ICMPForwarderResponder) Reject

func (r *ICMPForwarderResponder) Reject() error

Reject terminates the detached message and attempts to enqueue an administratively prohibited response without waiting for outbound capacity. It remains valid after replies and revalidates the forwarder and current destination policy. Once selected, the decision remains terminal on output error. Local output congestion may discard the response without error. It reports net.ErrClosed if the responder is already terminal or restricted to replies, including one returned by DetachForReplies, or if the originating forwarder is closed.

func (*ICMPForwarderResponder) Reply

func (r *ICMPForwarderResponder) Reply(payload []byte) error

Reply makes one immediate best-effort output attempt for a reverse ICMP message. Local output congestion may discard the message or any of its source fragments without error. Calls may be repeated or concurrent until a terminal action or while restricted to replies, with no ordering guarantee between concurrent calls. Any call may be retried after failure; each call revalidates the forwarder and current destination policy and copies payload before returning. It reports net.ErrClosed after a terminal action or when the originating forwarder is closed.

func (*ICMPForwarderResponder) ReplyEcho

func (r *ICMPForwarderResponder) ReplyEcho() error

ReplyEcho copies the detached Echo Request into an Echo Reply, preserving its identifier, sequence, and data. It reports syscall.EINVAL when the retained message is not an IPv4 or IPv6 Echo Request. Calls may be repeated or concurrent until a terminal action and may be followed by Drop or Reject. Its best-effort output behavior matches Reply. It reports net.ErrClosed if the responder is terminal or restricted to replies, including one returned by DetachForReplies, or if the originating forwarder is closed.

func (*ICMPForwarderResponder) ReplyIPPacket

func (r *ICMPForwarderResponder) ReplyIPPacket(packet []byte) error

ReplyIPPacket is the detached form of ICMPForwarderRequest.ReplyIPPacket. Calls may be repeated or concurrent until a terminal action or while restricted to replies. The packet is copied before return, concurrent calls have no ordering guarantee, and a failed call may be retried or followed by Drop or Reject while the responder remains active. Local output congestion may discard the packet or any of its source fragments without error. It reports net.ErrClosed after a terminal action or when the originating forwarder is closed.

func (*ICMPForwarderResponder) RestrictToReplies

func (r *ICMPForwarderResponder) RestrictToReplies() error

RestrictToReplies irreversibly discards the complete packet, message payload, and rejection quote while retaining Message metadata, Reply, ReplyIPPacket, and Done. On success, ReplyEcho, Reject, and Drop report net.ErrClosed. Calls made after a successful restriction or on a responder returned by DetachForReplies are no-ops that return nil. It reports net.ErrClosed only if the responder is terminal and does not itself count as a terminal action. The caller must invoke it while no other responder method is running. Previously returned snapshots remain valid and keep their storage live for as long as the caller retains them. After this method returns, Reply and ReplyIPPacket may again be called concurrently.

type ICMPMessage

type ICMPMessage struct {
	// Source is the source IP address.
	Source netip.Addr
	// Destination is the destination IP address.
	Destination netip.Addr
	// Type is the ICMP message type.
	Type uint8
	// Code is the subtype within Type.
	Code uint8
	// Body contains every byte after Type, Code, and Checksum. It must include
	// the message type's four-byte minimum body.
	Body []byte
}

ICMPMessage is the semantic representation of one ICMPv4 or ICMPv6 message. Source and Destination select the address family and, for ICMPv6, provide the checksum pseudo-header. IPPacket.ICMPMessage borrows Body from the packet; callers must replace or copy it before modifying unowned input. Validation and wire encoding treat IPv4-mapped IPv6 addresses as IPv4 without changing the receiver's address fields.

func (ICMPMessage) AppendBinary

func (m ICMPMessage) AppendBinary(dst []byte) ([]byte, error)

AppendBinary appends the complete ICMP message wire encoding to dst. Source and Destination select the family and contribute to the ICMPv6 pseudo-header checksum but are not themselves encoded. It validates every field before changing dst, does not retain any input slice, and permits the destination to share backing storage with Body. On validation failure it returns the original dst unchanged.

func (ICMPMessage) Echo

func (m ICMPMessage) Echo() (identifier, sequence uint16, payload []byte, ok bool)

Echo returns the identifier, sequence, and data of a complete family- appropriate Echo Request or Echo Reply. Payload aliases Body.

func (ICMPMessage) EchoReply

func (m ICMPMessage) EchoReply(source netip.Addr) (ICMPMessage, error)

EchoReply returns the semantic Echo Reply corresponding to m, using source as the reply source address. An explicit source is required because a request destination may be multicast, broadcast, or anycast and source selection requires IP routing state. The source argument must be an unzoned address in m's address family, but its ownership and address classification are not checked. Body aliases m.Body; MarshalBinary or AppendBinary calculates the reply checksum.

func (ICMPMessage) ICMPError

func (m ICMPMessage) ICMPError() (ICMPError, error)

ICMPError validates and decodes m as a supported ICMP error. QuotedPacket, QuotedPayload, and Extensions borrow m.Body, and available TCP or UDP ports are populated immediately. The parser accepts the intentionally truncated quotations that RFC 792 and RFC 4443 permit and does not apply socket- correlation policy.

func (ICMPMessage) IsEchoReply

func (m ICMPMessage) IsEchoReply() bool

IsEchoReply reports whether m is a complete family-appropriate Echo Reply. The four leading Body bytes hold the echo identifier and sequence.

func (ICMPMessage) IsEchoRequest

func (m ICMPMessage) IsEchoRequest() bool

IsEchoRequest reports whether m is a complete family-appropriate Echo Request. The four leading Body bytes hold the echo identifier and sequence.

func (ICMPMessage) IsError

func (m ICMPMessage) IsError() bool

IsError reports whether m has a supported, family-appropriate ICMP error type and code. It classifies the message without parsing its quoted packet; ICMPError performs complete quote validation.

func (ICMPMessage) MarshalBinary

func (m ICMPMessage) MarshalBinary() ([]byte, error)

MarshalBinary returns the complete ICMP message wire encoding. Source and Destination select the family and contribute to the ICMPv6 pseudo-header checksum but are not themselves encoded. MarshalBinary is semantically identical to AppendBinary(nil).

func (*ICMPMessage) SetEchoReply

func (m *ICMPMessage) SetEchoReply(identifier, sequence uint16, payload []byte) error

SetEchoReply replaces Type, Code, and Body with an Echo Reply containing identifier, sequence, and a copy of payload. Source and Destination must already select one valid address family. It leaves m unchanged on failure.

func (*ICMPMessage) SetEchoRequest

func (m *ICMPMessage) SetEchoRequest(identifier, sequence uint16, payload []byte) error

SetEchoRequest replaces Type, Code, and Body with an Echo Request containing identifier, sequence, and a copy of payload. Source and Destination must already select one valid address family. It leaves m unchanged on failure.

type ICMPv4Filter

type ICMPv4Filter struct {
	// contains filtered or unexported fields
}

ICMPv4Filter is a Linux-compatible ICMP_FILTER receive-type mask. A set bit blocks the corresponding representable ICMPv4 type, so the zero value accepts every type. Linux exposes 32 bits and x/net/ipv4 consequently uses the low five bits of method arguments; received types above 31 are outside the kernel mask and are always accepted.

func (*ICMPv4Filter) Accept

func (f *ICMPv4Filter) Accept(typ uint8)

Accept clears the mask bit selected by the low five bits of typ. Received ICMPv4 types above 31 remain unconditionally accepted by the socket.

func (*ICMPv4Filter) Block

func (f *ICMPv4Filter) Block(typ uint8)

Block sets the mask bit selected by the low five bits of typ. Received ICMPv4 types above 31 remain unconditionally accepted by the socket.

func (*ICMPv4Filter) SetAll

func (f *ICMPv4Filter) SetAll(block bool)

SetAll blocks every representable ICMPv4 type when block is true and accepts every type otherwise.

func (*ICMPv4Filter) WillBlock

func (f *ICMPv4Filter) WillBlock(typ uint8) bool

WillBlock reports the mask bit selected by the low five bits of typ. It may therefore report true for typ above 31 even though such received types are outside Linux's mask and remain accepted.

type ICMPv6Filter

type ICMPv6Filter struct {
	// contains filtered or unexported fields
}

ICMPv6Filter is an RFC 3542 ICMP6_FILTER receive-type mask. Its 256 bits cover every value of the ICMPv6 Type field. A set bit blocks that type, so the zero value accepts every type.

func (*ICMPv6Filter) Accept

func (f *ICMPv6Filter) Accept(typ uint8)

Accept permits packets whose ICMPv6 type has the supplied value.

func (*ICMPv6Filter) Block

func (f *ICMPv6Filter) Block(typ uint8)

Block rejects packets whose ICMPv6 type has the supplied value.

func (*ICMPv6Filter) SetAll

func (f *ICMPv6Filter) SetAll(block bool)

SetAll blocks every ICMPv6 type when block is true and accepts every type otherwise.

func (*ICMPv6Filter) WillBlock

func (f *ICMPv6Filter) WillBlock(typ uint8) bool

WillBlock reports whether packets with the supplied ICMPv6 type are rejected.

type IPConn

type IPConn struct {
	// contains filtered or unexported fields
}

IPConn is a connected or unconnected userspace IP protocol socket. It exchanges protocol payloads by default. SocketOptions.IPHeaderIncludedOnWrite and SocketOptions.IPHeaderIncludedOnRead independently expose complete packets on the write and read sides. ICMP receive filters and raw IPv6 checksum processing may be selected at creation and updated through IPConn methods.

func (*IPConn) Broadcast

func (c *IPConn) Broadcast() (bool, error)

Broadcast reports the raw socket's SO_BROADCAST-equivalent permission.

func (*IPConn) Close

func (c *IPConn) Close() error

Close unregisters the protocol socket and wakes blocked operations.

func (*IPConn) ConfirmPathMTU

func (c *IPConn) ConfirmPathMTU(mtu int) error

ConfirmPathMTU records application-level acknowledgement of a connected protocol probe. mtu is the complete IP packet size, not the payload size.

func (*IPConn) ConfirmPathMTUFor

func (c *IPConn) ConfirmPathMTUFor(target netip.Addr, mtu int) error

ConfirmPathMTUFor is the unconnected form of ConfirmPathMTU.

func (*IPConn) ExcludeSourceSpecificGroup

func (c *IPConn) ExcludeSourceSpecificGroup(group, source netip.Addr) error

ExcludeSourceSpecificGroup blocks source on an existing any-source raw membership. It returns EINVAL for an SSM group.

func (*IPConn) ICMPv4Filter

func (c *IPConn) ICMPv4Filter() (ICMPv4Filter, error)

ICMPv4Filter returns an independent snapshot of the socket's current receive-type filter.

func (*IPConn) ICMPv6Filter

func (c *IPConn) ICMPv6Filter() (ICMPv6Filter, error)

ICMPv6Filter returns an independent snapshot of the socket's current receive-type filter.

func (*IPConn) IPv6Checksum

func (c *IPConn) IPv6Checksum() (enabled bool, offset int, err error)

IPv6Checksum reports the checksum policy applied to IPv6 receive verification and ordinary payload writes. A disabled policy reports offset zero. ICMPv6 always reports enabled processing at offset 2.

func (*IPConn) IncludeSourceSpecificGroup

func (c *IPConn) IncludeSourceSpecificGroup(group, source netip.Addr) error

IncludeSourceSpecificGroup removes a source block from an existing any-source raw membership. It returns EINVAL for an SSM group.

func (*IPConn) Info

func (c *IPConn) Info() IPConnInfo

Info returns a diagnostic snapshot of the protocol socket and its receive queue.

func (*IPConn) JoinGroup

func (c *IPConn) JoinGroup(group netip.Addr) error

JoinGroup joins an any-source multicast group for a raw protocol socket. RFC 4604 SSM groups return EINVAL and require JoinSourceSpecificGroup.

func (*IPConn) JoinSourceSpecificGroup

func (c *IPConn) JoinSourceSpecificGroup(group, source netip.Addr) error

JoinSourceSpecificGroup adds source to an INCLUDE-mode raw membership, creating the membership when this is its first source.

func (*IPConn) LeaveGroup

func (c *IPConn) LeaveGroup(group netip.Addr) error

LeaveGroup leaves a raw protocol socket's multicast group.

func (*IPConn) LeaveSourceSpecificGroup

func (c *IPConn) LeaveSourceSpecificGroup(group, source netip.Addr) error

LeaveSourceSpecificGroup removes source from an INCLUDE-mode raw membership. It leaves the group when source was its final entry.

func (*IPConn) LocalAddr

func (c *IPConn) LocalAddr() net.Addr

LocalAddr returns the bound protocol address.

func (*IPConn) MulticastHopLimit

func (c *IPConn) MulticastHopLimit() (int, error)

MulticastHopLimit returns the raw socket's multicast TTL or Hop Limit.

func (*IPConn) MulticastLoopback

func (c *IPConn) MulticastLoopback() (bool, error)

MulticastLoopback reports whether raw multicast output is delivered locally.

func (*IPConn) MulticastSourceFilter

func (c *IPConn) MulticastSourceFilter(group netip.Addr) (MulticastSourceFilter, error)

MulticastSourceFilter returns a raw socket's complete source policy.

func (*IPConn) PathMTUDiscovery

func (c *IPConn) PathMTUDiscovery() (PathMTUDiscovery, error)

PathMTUDiscovery returns the Linux-compatible IP_MTU_DISCOVER policy used by subsequent protocol-payload writes.

func (*IPConn) Read

func (c *IPConn) Read(buffer []byte) (int, error)

Read receives from a connected remote endpoint.

func (*IPConn) ReadBatch

func (c *IPConn) ReadBatch(messages []SocketMessage, flags int) (int, error)

ReadBatch reads one or more IP protocol messages using the SocketMessage layout shared by x/net/ipv4 and x/net/ipv6. The first message follows the socket's blocking and deadline semantics; after it succeeds, the method drains only messages already queued. MessageFlagDontWait also makes the first read nonblocking.

func (*IPConn) ReadError

func (c *IPConn) ReadError() (*net.OpError, error)

ReadError returns the oldest queued asynchronous network error without blocking. An empty queue reports EAGAIN, like a Linux MSG_ERRQUEUE read on a nonblocking descriptor. Ordinary operations may consume the pending socket error without removing this entry from the extended error queue.

func (*IPConn) ReadFrom

func (c *IPConn) ReadFrom(buffer []byte) (int, net.Addr, error)

ReadFrom implements net.PacketConn.

func (*IPConn) ReadFromIP

func (c *IPConn) ReadFromIP(buffer []byte) (int, *net.IPAddr, error)

ReadFromIP acts like ReadFrom but returns an IPAddr.

func (*IPConn) ReadFromIPWithBuffer

func (c *IPConn) ReadFromIPWithBuffer(getBuffer func(sizeHint int) []byte) (int, *net.IPAddr, error)

ReadFromIPWithBuffer is the *net.IPAddr form of ReadFromWithBuffer.

It has the same callback, buffer ownership, truncation, and nil-callback semantics as ReadFromWithBuffer.

This is an experimental API and is not covered by the package's stability guarantees.

func (*IPConn) ReadFromWithBuffer

func (c *IPConn) ReadFromWithBuffer(getBuffer func(sizeHint int) []byte) (int, net.Addr, error)

ReadFromWithBuffer reads the next protocol payload like ReadFrom, obtaining the destination buffer lazily from getBuffer and returning its source address.

If a datagram is available, getBuffer is called once after it is dequeued. It is not called when the operation returns before a datagram is available because of an error or deadline. The callback receives the exposed datagram length as an advisory size hint: this is the protocol payload length, or the complete reassembled IP packet length when IPHeaderIncludedOnRead is enabled. It runs without c's connection state lock and must return promptly; it must not call a read method on c. The caller owns the returned slice, and c does not retain it. A nil callback returns EINVAL. A short returned slice truncates and consumes the datagram, matching ReadFrom.

This is an experimental API and is not covered by the package's stability guarantees.

func (*IPConn) ReadMsgIP

func (c *IPConn) ReadMsgIP(buffer, oob []byte) (n, oobn, flags int, address *net.IPAddr, err error)

ReadMsgIP reads one socket message and Linux-compatible packet info, hop-limit, and traffic-class ancillary data. An IPHeaderIncludedOnRead socket returns a complete reassembled IP packet instead of a protocol payload.

func (*IPConn) ReadWithBuffer

func (c *IPConn) ReadWithBuffer(getBuffer func(sizeHint int) []byte) (int, error)

ReadWithBuffer reads the next protocol payload from a connected remote endpoint like Read, obtaining the destination buffer lazily from getBuffer.

If a datagram is available, getBuffer is called once after it is dequeued. It is not called when the operation returns before a datagram is available because of an error or deadline. The callback receives the exposed datagram length as an advisory size hint: this is the protocol payload length, or the complete reassembled IP packet length when IPHeaderIncludedOnRead is enabled. It runs without c's connection state lock and must return promptly; it must not call a read method on c. The caller owns the returned slice, and c does not retain it. A nil callback returns EINVAL. A short returned slice truncates and consumes the datagram, matching Read.

This is an experimental API and is not covered by the package's stability guarantees.

func (*IPConn) ReceiveErrors

func (c *IPConn) ReceiveErrors() (bool, error)

ReceiveErrors reports whether asynchronous errors are retained for ReadError. When disabled, unconnected sockets do not report those errors; connected sockets report hard errors on reads or header-included writes. It also reports whether immediate failure to admit unicast or external-link non-unicast output is reported as ENOBUFS.

func (*IPConn) RemoteAddr

func (c *IPConn) RemoteAddr() net.Addr

RemoteAddr returns the connected peer, or nil for an unconnected socket.

func (*IPConn) SetBroadcast

func (c *IPConn) SetBroadcast(enabled bool) error

SetBroadcast changes the raw socket's SO_BROADCAST-equivalent permission.

func (*IPConn) SetDeadline

func (c *IPConn) SetDeadline(deadline time.Time) error

SetDeadline updates both read and write deadlines.

func (*IPConn) SetFlowLabel

func (c *IPConn) SetFlowLabel(label uint32) error

SetFlowLabel changes the default IPv6 Flow Label. Zero explicitly disables automatic labeling for this socket.

func (*IPConn) SetHopLimit

func (c *IPConn) SetHopLimit(hopLimit int) error

SetHopLimit changes the default IPv4 TTL or IPv6 Hop Limit for subsequent writes. Zero is valid only on a dedicated IPv6 socket; it is ambiguous on a dual-stack socket because IPv4 TTL zero is invalid. Per-packet message control data may override the value.

func (*IPConn) SetICMPv4Filter

func (c *IPConn) SetICMPv4Filter(filter ICMPv4Filter) error

SetICMPv4Filter atomically replaces the receive-type filter used by an IPv4 ICMP socket. Packets already queued for reading are not reconsidered.

func (*IPConn) SetICMPv6Filter

func (c *IPConn) SetICMPv6Filter(filter ICMPv6Filter) error

SetICMPv6Filter atomically replaces the receive-type filter used by an ICMPv6 socket. Packets already queued for reading are not reconsidered.

func (*IPConn) SetIPHeaderIncludedOnWrite

func (c *IPConn) SetIPHeaderIncludedOnWrite(enabled bool) error

SetIPHeaderIncludedOnWrite controls whether subsequent writes contain complete IP packets rather than protocol payloads.

func (*IPConn) SetIPv6Checksum

func (c *IPConn) SetIPv6Checksum(enabled bool, offset int) error

SetIPv6Checksum controls RFC 3542 checksum insertion and verification for ordinary upper-layer payloads on a non-ICMPv6 socket. When enabled, offset must be the even, non-negative byte offset of a 16-bit checksum field. When disabled, offset is ignored. Complete-packet writes remain caller-owned.

func (*IPConn) SetMulticastHopLimit

func (c *IPConn) SetMulticastHopLimit(hopLimit int) error

SetMulticastHopLimit changes the raw socket's multicast TTL or Hop Limit.

func (*IPConn) SetMulticastLoopback

func (c *IPConn) SetMulticastLoopback(enabled bool) error

SetMulticastLoopback controls raw multicast delivery to local memberships.

func (*IPConn) SetMulticastSourceFilter

func (c *IPConn) SetMulticastSourceFilter(group netip.Addr, filter MulticastSourceFilter) error

SetMulticastSourceFilter atomically replaces a raw socket's complete INCLUDE/EXCLUDE source policy for a previously joined group. An empty INCLUDE filter leaves that membership. EXCLUDE returns EINVAL for an RFC 4604 SSM group.

func (*IPConn) SetPathMTUDiscovery

func (c *IPConn) SetPathMTUDiscovery(mode PathMTUDiscovery) error

SetPathMTUDiscovery changes the Linux-compatible IP_MTU_DISCOVER policy for subsequent protocol-payload writes.

func (*IPConn) SetReadBuffer

func (c *IPConn) SetReadBuffer(bytes int) error

SetReadBuffer changes the approximate retained-memory capacity shared by the payload and asynchronous-error receive queues. Existing entries are retained when shrinking.

func (*IPConn) SetReadDeadline

func (c *IPConn) SetReadDeadline(deadline time.Time) error

SetReadDeadline updates pending and future reads.

func (*IPConn) SetReceiveErrors

func (c *IPConn) SetReceiveErrors(enabled bool) error

SetReceiveErrors controls whether asynchronous network errors are retained for ReadError. It also makes a write fail with ENOBUFS when immediate admission of unicast output or the external-link copy of multicast or broadcast output fails. It does not report packets displaced after admission. Receive-side non-unicast loopback copies remain best effort. By default, unconnected sockets do not report asynchronous ICMP errors, although correlated PMTU updates still apply. Connected sockets report hard errors on ordinary reads before queued payloads and on header-included writes; protocol-payload writes leave pending errors for reads. When enabled, eligible soft errors follow the same ordinary-operation policy and reach ReadError. Disabling the option clears the extended queue but preserves a pending ordinary error. When disabled, immediate output admission failures are silent.

func (*IPConn) SetTrafficClass

func (c *IPConn) SetTrafficClass(value int) error

SetTrafficClass changes the default IPv4 TOS or IPv6 Traffic Class byte.

func (*IPConn) SetWriteBuffer

func (c *IPConn) SetWriteBuffer(bytes int) error

SetWriteBuffer is a validated no-op because IP writes make one immediate bounded link-queue admission attempt and retain no per-socket transmit buffer to resize.

func (*IPConn) SetWriteDeadline

func (c *IPConn) SetWriteDeadline(deadline time.Time) error

SetWriteDeadline sets the deadline checked before future writes.

func (*IPConn) Write

func (c *IPConn) Write(payload []byte) (int, error)

Write sends one protocol payload or header-included packet to the connected endpoint.

func (*IPConn) WriteBatch

func (c *IPConn) WriteBatch(messages []SocketMessage, flags int) (int, error)

WriteBatch writes a prefix of IP payloads or header-included packets using scatter/gather buffers. MessageFlagDontWait is accepted for Linux compatibility; device admission is already nonblocking for every datagram write. Other flags are unsupported.

func (*IPConn) WriteMsgIP

func (c *IPConn) WriteMsgIP(payload, oob []byte, address *net.IPAddr) (n, oobn int, err error)

WriteMsgIP writes one protocol payload or header-included packet with Linux-compatible source, hop-limit, and traffic-class ancillary data. Like net.IPConn, it requires an unconnected socket and a non-nil destination.

func (*IPConn) WritePathMTUProbe

func (c *IPConn) WritePathMTUProbe(payload []byte) (int, error)

WritePathMTUProbe sends one connected protocol payload without IPv4 or IPv6 source fragmentation. The complete packet may exceed the confirmed PMTU but cannot exceed the first-hop MTU.

func (*IPConn) WritePathMTUProbeTo

func (c *IPConn) WritePathMTUProbeTo(payload []byte, target netip.Addr) (int, error)

WritePathMTUProbeTo is the unconnected netip form of WritePathMTUProbe.

func (*IPConn) WriteTo

func (c *IPConn) WriteTo(payload []byte, address net.Addr) (int, error)

WriteTo sends one protocol payload or header-included packet to an unconnected destination.

func (*IPConn) WriteToIP

func (c *IPConn) WriteToIP(payload []byte, address *net.IPAddr) (int, error)

WriteToIP acts like WriteTo but accepts an IPAddr directly.

type IPConnInfo

type IPConnInfo struct {
	// LocalAddress is the bound local address; an unspecified address denotes
	// a wildcard binding.
	LocalAddress netip.Addr
	// RemoteAddress is the connected peer, or an invalid address for an
	// unconnected socket.
	RemoteAddress netip.Addr
	// Protocol is the IANA IP protocol number carried by the socket.
	Protocol uint8
	// IPHeaderIncludedOnWrite reports whether writes contain complete IP
	// packets rather than protocol payloads.
	IPHeaderIncludedOnWrite bool
	// IPHeaderIncludedOnRead reports whether reads return complete reassembled
	// IP packets rather than protocol payloads.
	IPHeaderIncludedOnRead bool
	// Closed reports whether the socket was closed when the snapshot was taken.
	Closed bool
	// ReceiveQueuePackets is the number of complete messages awaiting a read.
	ReceiveQueuePackets int
	// ReceiveQueueBytes is the accounted payload and metadata retained by the
	// receive queue.
	ReceiveQueueBytes int
	// ReceiveQueueCapacity is the configured accounting-byte limit of the
	// combined payload and extended-error queues, not an exact heap limit.
	ReceiveQueueCapacity int
	// ReceiveErrors reports whether asynchronous network errors are retained
	// for ReadError. Otherwise, unconnected sockets do not report those errors;
	// connected sockets report hard errors on reads or header-included writes.
	// It also reports whether immediate failure to admit unicast or external-link
	// non-unicast output is reported as ENOBUFS.
	ReceiveErrors bool
	// ErrorQueueEntries is the number of extended errors awaiting ReadError.
	ErrorQueueEntries int
	// ErrorQueueBytes is the accounted metadata and quoted packet data retained
	// by the asynchronous error queue.
	ErrorQueueBytes int
	// ErrorsDropped counts extended-error entries discarded because the
	// configured receive-buffer budget was exhausted.
	ErrorsDropped uint64
	// PacketsSent counts successful IP socket write results. It includes writes
	// silently lost during bounded output admission under the default
	// ReceiveErrors policy and remains cumulative if bounded link scheduling
	// later drops a packet.
	PacketsSent uint64
	// BytesSent counts bytes represented by those successful writes.
	BytesSent uint64
	// PacketsReceived counts socket messages accepted into the receive queue.
	PacketsReceived uint64
	// BytesReceived counts bytes accepted in the configured read representation.
	BytesReceived uint64
	// PacketsDropped counts payloads rejected because the socket was closed or
	// its receive queue lacked capacity.
	PacketsDropped uint64
	// ICMPErrors counts matching asynchronous ICMP errors delivered to the
	// socket.
	ICMPErrors uint64
	// PathMTU is the complete-IP-packet PMTU for a connected unicast peer, or
	// zero when no such path exists.
	PathMTU int
	// PathMTUDiscovery is the Linux-compatible source-fragmentation and PMTU
	// policy used by subsequent writes.
	PathMTUDiscovery PathMTUDiscovery
	// HopLimit is the default unicast IPv4 TTL or IPv6 Hop Limit.
	HopLimit int
	// MulticastHopLimit is the default multicast IPv4 TTL or IPv6 Hop Limit.
	MulticastHopLimit int
	// MulticastLoopback reports whether transmitted multicast is delivered to
	// matching local memberships.
	MulticastLoopback bool
	// Broadcast reports whether IPv4 broadcast output is permitted.
	Broadcast bool
	// TrafficClass is the default IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass uint8
	// FlowLabel is the effective IPv6 Flow Label; it is zero for IPv4 sockets.
	FlowLabel uint32
	// LastError is the most recently recorded socket operation or asynchronous
	// network error.
	LastError error
}

IPConnInfo is a point-in-time diagnostic snapshot of one IP protocol socket. Traffic byte counters measure the representation exposed by the socket: protocol payloads by default or complete packets when the corresponding header option is enabled. Receive-queue byte values also include the stack's per-datagram accounting overhead.

type IPForwarder

type IPForwarder struct {
	// contains filtered or unexported fields
}

IPForwarder owns the single fallback handler for otherwise unhandled IP protocols installed on a stack. It may be installed before or after Stack.Start.

func NewIPForwarder

func NewIPForwarder(stack *Stack, options IPForwarderOptions, handler IPForwarderHandler) (*IPForwarder, error)

NewIPForwarder installs a fallback handler for otherwise unhandled IP protocols. A matching IPConn has priority, and TCP, UDP, ICMP, and IPv6 No Next Header never reach this handler. Only one IP forwarder may be active per stack. Promiscuous mode is required only for nonlocal destinations. Installing a forwarder does not start the stack.

func (*IPForwarder) Close

func (f *IPForwarder) Close() error

Close removes the otherwise unhandled IP protocol fallback handler and invalidates undecided requests. It does not wait for running handlers or replies.

func (IPForwarder) Done

func (f IPForwarder) Done() <-chan struct{}

Done is closed when the forwarder is closed directly or by Stack.Close.

func (*IPForwarder) Info

func (f *IPForwarder) Info() ForwarderInfo

Info returns one IP forwarder diagnostic snapshot.

type IPForwarderHandler

type IPForwarderHandler func(*IPForwarderRequest)

IPForwarderHandler processes one valid, reassembled upper-layer IP payload that matched neither a raw IP socket nor a built-in protocol. MIPS calls it synchronously with the same concurrency, blocking, ownership, and action rules as ICMPForwarderHandler. The handler must call Detach, DetachForReplies, Drop, Reject, or at least one Reply before returning.

type IPForwarderMessage

type IPForwarderMessage struct {
	// Source is the sender of the IP payload.
	Source netip.Addr
	// Destination is the original packet destination.
	Destination netip.Addr
	// Protocol is the IPv4 Protocol or final IPv6 Next Header value.
	Protocol uint8
	// HopLimit is the received IPv4 TTL or IPv6 Hop Limit.
	HopLimit uint8
	// TrafficClass is the received IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass uint8
	// FlowLabel is the received IPv6 Flow Label and is zero for IPv4.
	FlowLabel uint32
	// Payload contains the bytes following the IP or extension headers.
	Payload []byte
}

IPForwarderMessage describes one valid, reassembled upper-layer IP payload. Its ownership and lifetime are specified by the method that returned it.

type IPForwarderOptions

type IPForwarderOptions struct{}

IPForwarderOptions reserves otherwise unhandled IP protocol interception policy for future extension.

type IPForwarderRequest

type IPForwarderRequest struct {
	// contains filtered or unexported fields
}

IPForwarderRequest is one upper-layer IP payload not consumed by a raw IP socket or built-in protocol. The handler may reply repeatedly before selecting at most one terminal action.

func (*IPForwarderRequest) Detach

Detach transfers one IP request out of the synchronous handler lifetime and on success removes it from the forwarder's pending set. It returns a caller-owned message snapshot whose responder points back to the originating forwarder's state for output and Done; the forwarder does not retain the responder. The caller may discard it without a terminal action. Detach may be called after any number of Reply attempts; the responder remains available for further replies.

func (*IPForwarderRequest) DetachForReplies

func (r *IPForwarderRequest) DetachForReplies() (*IPForwarderResponder, error)

DetachForReplies transfers the IP request into an asynchronous responder that retains only Message metadata, Reply, and Done. Message().Payload is nil, and Reject and Drop report net.ErrClosed. Unlike Detach, it does not copy the upper-layer payload or retain an ICMP rejection quote. The caller owns the responder and may discard it without a terminal action. RestrictToReplies is an idempotent no-op on the returned responder.

func (*IPForwarderRequest) Drop

func (r *IPForwarderRequest) Drop() error

Drop consumes the IP payload without packet I/O and may follow any number of Reply attempts.

func (*IPForwarderRequest) Message

Message returns the validated upper-layer IP payload presented to the handler. Payload aliases packet-delivery storage, must not be modified, and is valid only until the handler returns. Message does not select an action.

func (*IPForwarderRequest) Reject

func (r *IPForwarderRequest) Reject() error

Reject consumes the IP payload and makes a best-effort attempt to enqueue the address-family protocol-unreachable response without waiting for outbound capacity. Local output congestion may discard the response without error. It reports syscall.EADDRNOTAVAIL when the intercepted destination is no longer admitted and syscall.ENETUNREACH when no return route remains. It may follow any number of Reply attempts.

func (*IPForwarderRequest) Reply

func (r *IPForwarderRequest) Reply(payload []byte) error

Reply sends one payload with the triggering protocol number from Destination to Source. Calls may be repeated or concurrent before a later terminal action and must finish before the handler returns. Each call uses the current Config.IP output defaults and makes one immediate best-effort output attempt. Local output congestion may discard the payload or any of its source fragments without error. Other errors may be retried.

type IPForwarderResponder

type IPForwarderResponder struct {
	// contains filtered or unexported fields
}

IPForwarderResponder owns one detached IP message. A responder returned by Detach owns its payload snapshot; DetachForReplies omits that snapshot. It may outlive the handler because the caller, not the forwarder, owns it. It retains access to the originating forwarder's state for output, diagnostics, and Done; neither the forwarder nor the stack retains the responder. Reply calls may be repeated while active or restricted to replies; Reject and Drop are available only while active. The responder may be discarded without a terminal action.

func (*IPForwarderResponder) Done

func (r *IPForwarderResponder) Done() <-chan struct{}

Done is closed when the originating protocol forwarder is closed directly or by Stack.Close.

func (*IPForwarderResponder) Drop

func (r *IPForwarderResponder) Drop() error

Drop terminates the detached input without packet I/O. It remains valid after replies while the responder is active. It reports net.ErrClosed when the responder is already terminal or restricted to replies, including one returned by DetachForReplies.

func (*IPForwarderResponder) Message

Message returns the detached IP metadata and independently owned payload. Payload is nil after DetachForReplies or RestrictToReplies. The caller may retain or modify a returned payload and must synchronize concurrent access.

func (*IPForwarderResponder) Reject

func (r *IPForwarderResponder) Reject() error

Reject terminates the detached payload and makes a best-effort attempt to enqueue a protocol-unreachable response. Local output congestion may discard the response without error. It remains valid after replies and revalidates current destination policy. It reports net.ErrClosed if the responder is already terminal or restricted to replies, including one returned by DetachForReplies, or if the originating forwarder is closed.

func (*IPForwarderResponder) Reply

func (r *IPForwarderResponder) Reply(payload []byte) error

Reply makes one immediate best-effort output attempt for a reverse protocol payload. Local output congestion may discard the payload or any of its source fragments without error. Calls may be repeated or concurrent until a terminal action or while restricted to replies, with no ordering guarantee between concurrent calls. Failed calls may be retried, and Drop or Reject may follow any number of replies while the responder remains active. Each call uses the current Config.IP output defaults. It reports net.ErrClosed after a terminal action or when the originating forwarder is closed.

func (*IPForwarderResponder) RestrictToReplies

func (r *IPForwarderResponder) RestrictToReplies() error

RestrictToReplies irreversibly discards the upper-layer payload and rejection quote while retaining Message metadata, Reply, and Done. On success, Reject and Drop report net.ErrClosed. Calls made after a successful restriction or on a responder returned by DetachForReplies are no-ops that return nil. It reports net.ErrClosed only if the responder is terminal and does not itself count as a terminal action. The caller must invoke it while no other responder method is running. A previously returned Message().Payload slice remains valid and keeps its storage live for as long as the caller retains it. After this method returns, Reply may again be called concurrently.

type IPPacket

type IPPacket struct {
	// Source is the packet's source address.
	Source netip.Addr
	// Destination is the packet's destination address.
	Destination netip.Addr
	// Protocol is the IPv4 Protocol or IPv6 base-header Next Header value.
	Protocol int
	// HopLimit is the IPv4 TTL or IPv6 Hop Limit, including an explicit zero.
	HopLimit int
	// TrafficClass is the complete IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass int
	// FlowLabel is the IPv6 Flow Label and must be zero for IPv4.
	FlowLabel uint32
	// Identification is the IPv4 Identification field and must be zero for IPv6.
	Identification uint16
	// DontFragment is the IPv4 Don't Fragment flag and must be false for IPv6.
	DontFragment bool
	// MoreFragments is the IPv4 More Fragments flag and must be false for IPv6.
	MoreFragments bool
	// FragmentOffset is the IPv4 fragment offset in bytes and must be zero for
	// IPv6. A nonzero value must be aligned to eight bytes.
	FragmentOffset int
	// IPv4Options contains the exact IPv4 option area, including received padding.
	// Construction also accepts an unpadded option sequence. AppendBinary adds
	// final alignment and normalizes every byte after End to zero;
	// AppendRawBinary preserves supplied bytes and only adds alignment padding.
	IPv4Options []byte
	// Payload is the complete IP payload. It includes IPv6 extension headers.
	Payload []byte
}

IPPacket is the semantic representation of one complete wire IPv4 or IPv6 packet. A value may therefore be an IPv4 fragment or contain an IPv6 Fragment header; it does not imply that the original datagram is complete. For IPv6, Protocol is the base header's immediate Next Header value and Payload includes any extension headers. ParseIPPacket borrows IPv4Options and Payload from its input; callers must replace or copy those slices before modifying data they do not own. Construction normalizes IPv4-mapped IPv6 addresses to IPv4.

func ParseIPPacket

func ParseIPPacket(packet []byte) (IPPacket, error)

ParseIPPacket validates packet and returns a zero-copy semantic value. It validates declared lengths, the IPv4 header checksum, option framing, IPv6 extension framing, fragment fields, and extension placement. It accepts one complete wire fragment but performs no stateful reassembly. Bytes beyond the declared IP length are ignored as link-layer padding.

func (IPPacket) AppendBinary

func (p IPPacket) AppendBinary(dst []byte) ([]byte, error)

AppendBinary appends the complete packet wire encoding to dst. It validates every field and the extension chain before changing dst and does not retain any input slice. The destination may share backing storage with IPv4Options or Payload. It calculates the IPv4 header checksum but never calculates an upper-layer checksum; callers must encode any TCP, UDP, or ICMP value before assigning it to Payload. On validation failure it returns the original dst unchanged.

func (IPPacket) AppendRawBinary

func (p IPPacket) AppendRawBinary(dst []byte) ([]byte, error)

AppendRawBinary appends a packet with a valid IPv4 or IPv6 fixed header while treating IPv4Options and the IPv6 Protocol and Payload fields as opaque wire data. For IPv4, it also permits malformed option framing and fragment payload sizes or reconstructed bounds that violate fragment semantics, while still requiring an exactly representable fragment offset. Unlike AppendBinary, it preserves every supplied option and payload byte, including IPv4 End padding and IPv6 extension reserved fields and PadN contents. It calculates the IPv4 header checksum but no upper-layer checksum, does not retain input storage, and supports output overlapping Payload or IPv4Options. On failure it returns dst unchanged.

func (IPPacket) Fragment

func (p IPPacket) Fragment() (IPPacketFragmentView, bool)

Fragment returns the fragment metadata carried by p. For a valid packet it returns false only when an IPv4 packet is unfragmented or an IPv6 packet has no Fragment header. The returned payload borrows p.Payload.

func (IPPacket) ICMPMessage

func (p IPPacket) ICMPMessage() (ICMPMessage, error)

ICMPMessage validates and decodes the packet's family-appropriate ICMP upper layer.

func (IPPacket) IPv4HeaderOptions

func (p IPPacket) IPv4HeaderOptions() ([]IPv4HeaderOption, error)

IPv4HeaderOptions parses the exact IPv4 option sequence. The returned slice owns its descriptors, but each Data field borrows IPv4Options. End is returned as the final descriptor; its following received padding is omitted.

func (IPPacket) IPv6ExtensionHeaders

func (p IPPacket) IPv6ExtensionHeaders() (headers []IPv6ExtensionHeader, protocol int, payload []byte, err error)

IPv6ExtensionHeaders returns the structurally traversable extension-header chain, the final Next Header value, and the remaining payload. The returned descriptor slice is caller-owned; header Data and payload borrow p.Payload. ESP, No Next Header, and unknown values terminate traversal. Bytes following No Next Header are returned as payload for lossless structural reconstruction even though UpperLayer intentionally ignores them. A non-atomic Fragment is the final descriptor: protocol is its Next Header and payload is its raw fragmentable data, which is not traversed as a complete extension chain.

func (IPPacket) MarshalBinary

func (p IPPacket) MarshalBinary() ([]byte, error)

MarshalBinary returns the complete packet wire encoding. It is semantically identical to AppendBinary(nil).

func (IPPacket) MarshalFragments

func (p IPPacket) MarshalFragments(mtu int, ipv6Identification uint32) ([][]byte, error)

MarshalFragments returns one or more owned wire packets whose complete L3 length is no larger than mtu. A fitting packet produces the same encoding as MarshalBinary. IPv4 uses the packet's Identification, honors DontFragment, and may refragment an existing fragment while preserving its original offset and More Fragments state. IPv6 uses ipv6Identification only when fragmentation is required, replaces an atomic Fragment header, and refuses to refragment a non-atomic fragment. For a newly fragmented datagram, the caller is responsible for satisfying the Identification non-reuse requirement for the IPv4 source, destination, and protocol tuple or the IPv6 source and final-destination pair. IPv6 returns syscall.EMSGSIZE if mtu cannot carry the complete known header chain in the first fragment as required by RFC 7112. The IPv6 identification argument is ignored for IPv4 and fitting packets. No partial result is returned on failure.

func (IPPacket) MarshalRawBinary

func (p IPPacket) MarshalRawBinary() ([]byte, error)

MarshalRawBinary returns the complete packet wire encoding produced by AppendRawBinary(nil).

func (*IPPacket) SetIPv4HeaderOptions

func (p *IPPacket) SetIPv4HeaderOptions(options []IPv4HeaderOption) error

SetIPv4HeaderOptions replaces IPv4Options with the encoded option sequence. It preserves unknown types, duplicates, and order, copies every input Data slice, and leaves p unchanged on failure. End must be last. MarshalBinary or AppendBinary adds any final four-byte header alignment.

func (*IPPacket) SetIPv6ExtensionHeaders

func (p *IPPacket) SetIPv6ExtensionHeaders(headers []IPv6ExtensionHeader, protocol int, payload []byte) error

SetIPv6ExtensionHeaders replaces Protocol and Payload with one complete extension chain and final upper-layer payload. It generates every Next Header link, copies all caller storage, and leaves p unchanged on failure. Each header Data excludes its Next Header byte and must otherwise contain a complete header-specific wire representation. A non-atomic Fragment must be the final descriptor; only in that case may protocol itself identify a traversable extension header whose bytes begin in payload.

func (*IPPacket) SetRawIPv6ExtensionHeaders

func (p *IPPacket) SetRawIPv6ExtensionHeaders(headers []IPv6ExtensionHeader, protocol int, payload []byte) error

SetRawIPv6ExtensionHeaders replaces Protocol and Payload with an explicitly ordered extension chain and final payload. It links recognized extension header descriptors and copies all caller storage, but deliberately does not validate header framing, ordering, duplication, options, or Fragment semantics. AppendRawBinary preserves the resulting bytes; AppendBinary still applies the ordinary validation and sender normalization rules. The method leaves p unchanged on failure.

func (IPPacket) TCPSegment

func (p IPPacket) TCPSegment() (TCPSegment, error)

TCPSegment validates and decodes the packet's TCP upper layer. The three unexposed reserved bits are accepted and intentionally omitted from the semantic result, as Linux does; encoding writes them as zero. The historic NS bit remains available through TCPFlagNS for wire compatibility.

func (IPPacket) UDPDatagram

func (p IPPacket) UDPDatagram() (UDPDatagram, error)

UDPDatagram validates and decodes the packet's UDP upper layer. A shorter UDP Length trims IP-layer padding; a length larger than the upper layer is invalid.

func (IPPacket) UpperLayer

func (p IPPacket) UpperLayer() (protocol int, payload []byte, err error)

UpperLayer returns the final IPv4 protocol or IPv6 Next Header value and the corresponding upper-layer bytes. The returned slice aliases Payload. For IPv6 it walks supported extension headers without applying Stack routing or option-admission policy. It rejects a non-atomic fragment because that packet does not contain a complete upper-layer unit.

type IPPacketFragmentView

type IPPacketFragmentView struct {
	// Protocol identifies the fragmentable payload's first header.
	Protocol int
	// Identification is the IPv4 or IPv6 fragment identification value.
	Identification uint32
	// Offset is the fragment's byte offset in the original fragmentable part.
	Offset int
	// MoreFragments reports whether another fragment follows this one.
	MoreFragments bool
	// Payload is the fragmentable data carried by this packet.
	Payload []byte
}

IPPacketFragmentView describes the fragment carried by one IPPacket. Payload borrows the packet's payload storage. For IPv6, Protocol is the Fragment header's Next Header value and Payload begins immediately after that header. The scalar fields are snapshots and do not mutate the packet.

func (IPPacketFragmentView) IsAtomic

func (f IPPacketFragmentView) IsAtomic() bool

IsAtomic reports whether the view describes an IPv6 atomic fragment. A view returned for IPv4 is always non-atomic because an unfragmented IPv4 packet has no fragment header to expose.

type IPPacketReassembly

type IPPacketReassembly struct {
	// contains filtered or unexported fields
}

IPPacketReassembly incrementally reassembles one fragmented IP packet. The zero value is ready for use. Add copies every retained byte, so callers may reuse the input packet's backing storage after Add returns.

An IPPacketReassembly must not be copied after its first use. Its methods are not safe for concurrent use; callers processing one datagram concurrently must serialize Add and Reset calls.

func (*IPPacketReassembly) Add

func (r *IPPacketReassembly) Add(fragment IPPacket) (packet IPPacket, complete bool, err error)

Add adds one non-atomic IPv4 or IPv6 fragment. When complete is false and err is nil, the receiver remains incomplete; the input was either retained or recognized as a duplicate. A completed packet owns its IPv4Options and Payload storage, and the receiver is reset for reuse before Add returns.

An unfragmented packet, an IPv6 atomic fragment, an invalid standalone packet, or a fragment belonging to a different packet reports syscall.EINVAL without changing an existing reassembly. Other standalone validation errors are returned without changing it. A range already fully covered by retained fragments is a duplicate: its payload, ECN, and header metadata are ignored, but a duplicate final fragment may establish the packet's final length. A duplicate never completes reassembly by itself. Once a valid fragment is associated with the current packet, a conflicting length, ECN value, or partial overlap invalidates and resets the entire in-progress reassembly.

func (*IPPacketReassembly) Reset

func (r *IPPacketReassembly) Reset()

Reset discards an incomplete packet and releases all retained storage. Reset is idempotent.

type IPSocketDefaults

type IPSocketDefaults struct {
	DatagramSocketDefaults

	// IPHeaderIncludedOnWrite makes new IPConn writes contain a complete IPv4
	// or IPv6 packet instead of only the upper-layer protocol payload.
	IPHeaderIncludedOnWrite bool
	// IPHeaderIncludedOnRead makes new IPConn reads return the complete,
	// reassembled IP packet instead of only the upper-layer protocol payload.
	IPHeaderIncludedOnRead bool
}

IPSocketDefaults configures policies inherited by newly created IP protocol sockets. Zero fields retain the package defaults.

type IPv4ControlMessage

type IPv4ControlMessage struct {
	// TTL is the received or requested time to live. A received value may be
	// zero; zero selects the output default when marshaling.
	TTL int
	// TOS is the received or requested type-of-service byte. Zero selects the
	// output default when marshaling.
	TOS int
	// Src selects a managed source address when marshaling.
	Src netip.Addr
	// Dst is the IP header destination populated by Parse and is not marshaled.
	Dst netip.Addr
	// IfIndex is the embedding-link index. MIPS supports only zero.
	IfIndex int
}

IPv4ControlMessage represents per-packet IPv4 metadata carried by UDPConn.ReadMsgUDP, UDPConn.WriteMsgUDP, IPConn.ReadMsgIP, and IPConn.WriteMsgIP. Src is used when sending and Dst is populated from the received IP header destination. MIPS has one embedding link, so IfIndex is always zero and a nonzero value cannot be marshaled.

func (*IPv4ControlMessage) Marshal

func (message *IPv4ControlMessage) Marshal() ([]byte, error)

Marshal returns the fixed Linux 64-bit little-endian ancillary encoding used by MIPS on every host. Zero-valued fields are omitted and select stack defaults.

func (*IPv4ControlMessage) Parse

func (message *IPv4ControlMessage) Parse(control []byte) error

Parse replaces message with metadata decoded from the fixed ancillary encoding returned by MIPS message reads.

type IPv4HeaderOption

type IPv4HeaderOption struct {
	// Type is the complete eight-bit IPv4 option type.
	Type uint8
	// Data is the option value following Type and Length.
	Data []byte
}

IPv4HeaderOption is one IPv4 header option in semantic wire order. Data excludes the Type and Length bytes. End and NOP require empty Data; every other type is encoded with a Length byte. IPv4HeaderOptions returns Data slices that borrow IPPacket.IPv4Options, while SetIPv4HeaderOptions copies every Data slice.

func (IPv4HeaderOption) Class

func (o IPv4HeaderOption) Class() uint8

Class returns the IPv4 option's two-bit class field.

func (IPv4HeaderOption) Copied

func (o IPv4HeaderOption) Copied() bool

Copied reports the IPv4 option's copied flag, which requests copying into every fragment.

func (IPv4HeaderOption) Number

func (o IPv4HeaderOption) Number() uint8

Number returns the IPv4 option's five-bit option number.

func (IPv4HeaderOption) RouterAlert

func (o IPv4HeaderOption) RouterAlert() (uint16, bool)

RouterAlert returns the complete two-byte Router Alert value when o is a well-formed RFC 2113 option. The standalone codec preserves unassigned values; Stack protocol handling currently recognizes only value zero.

func (*IPv4HeaderOption) SetRouterAlert

func (o *IPv4HeaderOption) SetRouterAlert(value uint16)

SetRouterAlert replaces o with a Router Alert option containing value.

type IPv6ControlMessage

type IPv6ControlMessage struct {
	// TrafficClass is the received or requested traffic-class byte. Zero
	// selects the output default when marshaling.
	TrafficClass int
	// HopLimit is the received or requested hop limit. A received value may be
	// zero; zero selects the output default when marshaling.
	HopLimit int
	// FlowLabel is the received or requested 20-bit IPv6 Flow Label. Zero
	// selects automatic labeling when marshaling through this structured API.
	FlowLabel uint32
	// Src selects a managed source address when marshaling.
	Src netip.Addr
	// Dst is the IP header destination populated by Parse and is not marshaled.
	Dst netip.Addr
	// IfIndex is the embedding-link index. MIPS supports only zero.
	IfIndex int
}

IPv6ControlMessage represents per-packet IPv6 metadata carried by UDPConn.ReadMsgUDP, UDPConn.WriteMsgUDP, IPConn.ReadMsgIP, and IPConn.WriteMsgIP. Src is used when sending and Dst is populated when parsing received control data. MIPS has one embedding link, so IfIndex is always zero and a nonzero value cannot be marshaled.

func (*IPv6ControlMessage) Marshal

func (message *IPv6ControlMessage) Marshal() ([]byte, error)

Marshal returns the fixed Linux 64-bit little-endian ancillary encoding used by MIPS on every host. Zero-valued fields are omitted and select stack defaults.

func (*IPv6ControlMessage) Parse

func (message *IPv6ControlMessage) Parse(control []byte) error

Parse replaces message with metadata decoded from the fixed ancillary encoding returned by MIPS message reads.

type IPv6ExtensionHeader

type IPv6ExtensionHeader struct {
	// Type identifies the extension-header wire format.
	Type uint8
	// Data contains the raw header bytes following Next Header.
	Data []byte
}

IPv6ExtensionHeader is one structurally traversable IPv6 extension header. Type is the value naming the header in the preceding Next Header field. Data excludes this header's own Next Header byte but includes every remaining raw field, including its length byte when present. IPv6ExtensionHeaders returns Data slices that borrow IPPacket.Payload. Parsing preserves received PadN data and sender-reserved fields; IPPacket.MarshalBinary and IPPacket.AppendBinary clear PadN data and the reserved Fragment fields. Authentication and Mobility data remains opaque because changing it would invalidate the header's ICV or checksum.

func (IPv6ExtensionHeader) Fragment

func (h IPv6ExtensionHeader) Fragment() (offset int, moreFragments bool, identification uint32, ok bool)

Fragment decodes an IPv6 Fragment header. Offset is returned in bytes. Reserved fields are ignored as required for reception by RFC 8200. The method returns false when h is not an exactly sized Fragment header.

func (IPv6ExtensionHeader) Options

Options parses a Hop-by-Hop or Destination Options header. The returned slice owns its descriptors, but every Data field borrows h.Data. Unknown option types, action bits, duplicates, and padding remain in wire order.

func (*IPv6ExtensionHeader) SetFragment

func (h *IPv6ExtensionHeader) SetFragment(offset int, moreFragments bool, identification uint32) error

SetFragment replaces h with an IPv6 Fragment header. Offset is measured in bytes and must be in 0..65528 and aligned to eight bytes. Reserved fields are encoded as zero. The method leaves h unchanged on failure.

func (*IPv6ExtensionHeader) SetOptions

func (h *IPv6ExtensionHeader) SetOptions(options []IPv6ExtensionOption) error

SetOptions replaces Data in a Hop-by-Hop or Destination Options header. It preserves unknown types, duplicates, and order, copies every input Data slice, and adds canonical trailing Pad1 or PadN alignment. It leaves h unchanged on failure.

type IPv6ExtensionOption

type IPv6ExtensionOption struct {
	// Type is the complete eight-bit IPv6 option type.
	Type uint8
	// Data is the option value following Type and Opt Data Len.
	Data []byte
}

IPv6ExtensionOption is one option in a Hop-by-Hop or Destination Options header. Data excludes Option Type and Opt Data Len. Options returns Data slices that borrow IPv6ExtensionHeader.Data, while SetOptions copies them.

func (IPv6ExtensionOption) Action

func (o IPv6ExtensionOption) Action() uint8

Action returns the option's two-bit RFC 8200 action on an unrecognized type.

func (IPv6ExtensionOption) MayChangeInTransit

func (o IPv6ExtensionOption) MayChangeInTransit() bool

MayChangeInTransit reports the RFC 8200 mutable-data flag.

func (IPv6ExtensionOption) RouterAlert

func (o IPv6ExtensionOption) RouterAlert() (uint16, bool)

RouterAlert returns the complete two-byte Router Alert value when o is a well-formed RFC 2711 option. The standalone codec preserves unassigned values; Stack protocol handling currently recognizes only value zero.

func (*IPv6ExtensionOption) SetRouterAlert

func (o *IPv6ExtensionOption) SetRouterAlert(value uint16)

SetRouterAlert replaces o with a Router Alert option containing value.

type KeepAliveConfig

type KeepAliveConfig struct {
	// Idle is the inactivity interval before the first probe.
	Idle time.Duration
	// Interval is the delay between unanswered probes.
	Interval time.Duration
	// Count is the number of unanswered probes allowed before failure.
	Count int
}

KeepAliveConfig configures TCP keepalive probing. Every field must be positive when supplied to SetKeepAliveConfig.

type ListenConfig

type ListenConfig struct {
	// Options contains creation policies applied before binding. The call reads
	// but does not retain the slice. Repeated option kinds use the last value.
	Options []SocketOption
}

ListenConfig contains options applied before a socket is bound. The zero value is ready for use.

func (*ListenConfig) ListenIP

func (config *ListenConfig) ListenIP(ctx context.Context, stack *Stack, network string, local netip.Addr) (net.PacketConn, error)

ListenIP binds an unconnected IP protocol socket on stack. The returned net.PacketConn has dynamic type *IPConn.

func (*ListenConfig) ListenTCP

func (config *ListenConfig) ListenTCP(ctx context.Context, stack *Stack, network string, local netip.AddrPort) (net.Listener, error)

ListenTCP binds a TCP listener on stack. The returned net.Listener has dynamic type *TCPListener.

func (*ListenConfig) ListenUDP

func (config *ListenConfig) ListenUDP(ctx context.Context, stack *Stack, network string, local netip.AddrPort) (net.PacketConn, error)

ListenUDP binds an unconnected UDP packet socket on stack. The returned net.PacketConn has dynamic type *UDPConn.

type MulticastSourceFilter

type MulticastSourceFilter struct {
	// Mode selects an INCLUDE or EXCLUDE source set.
	Mode MulticastSourceFilterMode
	// Sources is copied by SetMulticastSourceFilter and may be reused by the
	// caller as soon as that method returns.
	Sources []netip.Addr
}

MulticastSourceFilter is the complete per-socket RFC 3678 filter state for one group. SetMulticastSourceFilter copies Sources before returning, and MulticastSourceFilter returns an independent, address-sorted snapshot.

type MulticastSourceFilterMode

type MulticastSourceFilterMode uint8

MulticastSourceFilterMode identifies an RFC 3678 full-state source policy.

const (
	// MulticastSourceFilterExclude receives packets from every source except
	// those listed. An empty EXCLUDE list is an any-source membership. Its
	// value matches Linux and RFC 3678's MCAST_EXCLUDE.
	MulticastSourceFilterExclude MulticastSourceFilterMode = iota
	// MulticastSourceFilterInclude receives packets only from listed sources.
	// An empty INCLUDE list is equivalent to leaving the group. Its value
	// matches Linux and RFC 3678's MCAST_INCLUDE.
	MulticastSourceFilterInclude
)

type PathMTUDiscovery

type PathMTUDiscovery int

PathMTUDiscovery is the Linux-compatible IP_MTU_DISCOVER policy used by UDP and IP protocol sockets. Its numeric values match IP_PMTUDISC_* so callers translating a Linux socket policy do not need another mapping table.

const (
	// PathMTUDiscoveryDont permits source fragmentation and leaves IPv4 DF
	// clear. Validated destination PMTU information still selects the local
	// fragment size, matching Linux IP_PMTUDISC_DONT route-cache behavior.
	PathMTUDiscoveryDont PathMTUDiscovery = iota
	// PathMTUDiscoveryWant uses the destination PMTU, sets IPv4 DF when a
	// datagram fits, and source-fragments it when needed.
	PathMTUDiscoveryWant
	// PathMTUDiscoveryDo uses the destination PMTU, always sets IPv4 DF, and
	// reports EMSGSIZE instead of source-fragmenting an oversized datagram.
	PathMTUDiscoveryDo
	// PathMTUDiscoveryProbe ignores the destination PMTU, uses the embedding
	// interface MTU, sets IPv4 DF, and reports EMSGSIZE above that MTU.
	PathMTUDiscoveryProbe
	// PathMTUDiscoveryInterface ignores destination PMTU updates, uses the
	// embedding interface MTU, leaves IPv4 DF clear, and reports EMSGSIZE above
	// that MTU.
	PathMTUDiscoveryInterface
	// PathMTUDiscoveryOmit ignores destination PMTU updates, uses the embedding
	// interface MTU, leaves IPv4 DF clear, and permits source fragmentation.
	PathMTUDiscoveryOmit
)

type RXChecksumOffload

type RXChecksumOffload struct {
	// contains filtered or unexported fields
}

RXChecksumOffload selects checksum verification delegated to the input link. Its zero value retains all verification. An enabled category requires the caller to supply valid complete packets with correct checksums or an equivalent validation guarantee.

Offload does not disable framing or protocol checks, including IPv6 UDP's zero-checksum prohibition. Reassembled payloads, public codecs, checksums configured for other raw IPv6 protocols, and independent ICMP extension checksums retain their verification. Configure the policy before delivering input packets.

func (RXChecksumOffload) ICMPv4

func (o RXChecksumOffload) ICMPv4() bool

ICMPv4 reports whether outer ICMPv4 checksum offload is enabled.

func (RXChecksumOffload) ICMPv6

func (o RXChecksumOffload) ICMPv6() bool

ICMPv6 reports whether ICMPv6 checksum offload is enabled.

func (RXChecksumOffload) IGMP

func (o RXChecksumOffload) IGMP() bool

IGMP reports whether IGMP checksum offload is enabled.

func (RXChecksumOffload) IPv4Header

func (o RXChecksumOffload) IPv4Header() bool

IPv4Header reports whether IPv4 header checksum offload is enabled.

func (*RXChecksumOffload) SetICMPv4

func (o *RXChecksumOffload) SetICMPv4(enabled bool) *RXChecksumOffload

SetICMPv4 controls outer ICMPv4 checksum offload and returns o for chained configuration. Independent extension checksums remain verified.

func (*RXChecksumOffload) SetICMPv6

func (o *RXChecksumOffload) SetICMPv6(enabled bool) *RXChecksumOffload

SetICMPv6 controls ICMPv6 checksum offload, including MLD, and returns o for chained configuration. Independent extension checksums remain verified.

func (*RXChecksumOffload) SetIGMP

func (o *RXChecksumOffload) SetIGMP(enabled bool) *RXChecksumOffload

SetIGMP controls IGMP checksum offload and returns o for chained configuration.

func (*RXChecksumOffload) SetIPv4Header

func (o *RXChecksumOffload) SetIPv4Header(enabled bool) *RXChecksumOffload

SetIPv4Header controls IPv4 header checksum offload, including fragment headers, and returns o for chained configuration.

func (*RXChecksumOffload) SetTCP

func (o *RXChecksumOffload) SetTCP(enabled bool) *RXChecksumOffload

SetTCP controls TCP checksum offload for both address families and returns o for chained configuration.

func (*RXChecksumOffload) SetUDP

func (o *RXChecksumOffload) SetUDP(enabled bool) *RXChecksumOffload

SetUDP controls UDP checksum offload for both address families and returns o for chained configuration. IPv6 zero-checksum datagrams remain invalid.

func (RXChecksumOffload) TCP

func (o RXChecksumOffload) TCP() bool

TCP reports whether TCP checksum offload is enabled.

func (RXChecksumOffload) UDP

func (o RXChecksumOffload) UDP() bool

UDP reports whether UDP checksum offload is enabled.

type Route

type Route struct {
	// Destination is the admitted unicast prefix.
	Destination netip.Prefix
	// Source pins the preferred local source address when valid.
	Source netip.Addr
	// Metric breaks ties between routes with equal prefix length.
	Metric uint32
}

Route admits one unicast destination prefix. Source optionally pins the preferred local source address; Metric breaks ties between equal prefixes. A source-less route may also carry transparent forwarder output when Promiscuous is enabled without a same-family local address. The embedding link remains responsible for next-hop selection.

type SocketErrorControlMessage

type SocketErrorControlMessage struct {
	// Errno is the Linux errno number associated with the failed operation.
	Errno uint32
	// Origin identifies the subsystem that generated the error.
	Origin SocketErrorOrigin
	// Type is the ICMP or ICMPv6 classification.
	Type uint8
	// Code is the ICMP or ICMPv6 subtype associated with Type.
	Code uint8
	// Info contains the discovered MTU or parameter-problem pointer when the
	// ICMP type defines one.
	Info uint32
	// Data is the protocol-specific sock_extended_err data field.
	Data uint32
	// Offender is the router or destination that reported the failure. It must
	// be a valid, unzoned IPv4 or IPv6 address when marshaling. An IPv4-mapped
	// IPv6 address retains its IPv6 representation.
	Offender netip.Addr
}

SocketErrorControlMessage represents the structured fields of one Linux sock_extended_err ancillary record used by MessageFlagErrorQueue reads.

func (SocketErrorControlMessage) AppendBinary

func (message SocketErrorControlMessage) AppendBinary(dst []byte) ([]byte, error)

AppendBinary appends one complete, aligned Linux 64-bit little-endian ancillary record to dst. Offender selects IP_RECVERR or IPV6_RECVERR; IPv4-mapped addresses retain their IPv6 sockaddr representation. Reserved fields and sockaddr fields not represented by SocketErrorControlMessage are encoded as zero. Numeric fields, including unrecognized Origin, Type, and Code values, are encoded unchanged without applying ICMP policy. On validation failure it returns the original dst unchanged.

func (SocketErrorControlMessage) MarshalBinary

func (message SocketErrorControlMessage) MarshalBinary() ([]byte, error)

MarshalBinary returns one complete, aligned Linux 64-bit little-endian ancillary record. It is semantically identical to AppendBinary(nil).

func (*SocketErrorControlMessage) Parse

func (message *SocketErrorControlMessage) Parse(control []byte) error

Parse replaces message with the one error record found in control. Other well-formed ancillary records are ignored so packet metadata may coexist in the same OOB buffer. Records with an AF_UNSPEC offender are rejected because SocketErrorControlMessage does not carry the cmsg address family separately.

type SocketErrorOrigin

type SocketErrorOrigin uint8

SocketErrorOrigin identifies the Linux sock_extended_err producer encoded in a SocketErrorControlMessage.

const (
	// SocketErrorOriginNone is Linux SO_EE_ORIGIN_NONE.
	SocketErrorOriginNone SocketErrorOrigin = iota
	// SocketErrorOriginLocal is Linux SO_EE_ORIGIN_LOCAL.
	SocketErrorOriginLocal
	// SocketErrorOriginICMP is Linux SO_EE_ORIGIN_ICMP.
	SocketErrorOriginICMP
	// SocketErrorOriginICMP6 is Linux SO_EE_ORIGIN_ICMP6.
	SocketErrorOriginICMP6
	// SocketErrorOriginTXStatus is Linux SO_EE_ORIGIN_TXSTATUS.
	SocketErrorOriginTXStatus
	// SocketErrorOriginZeroCopy is Linux SO_EE_ORIGIN_ZEROCOPY.
	SocketErrorOriginZeroCopy
	// SocketErrorOriginTXTime is Linux SO_EE_ORIGIN_TXTIME.
	SocketErrorOriginTXTime
)

type SocketMessage

type SocketMessage struct {
	// Buffers contains the contiguous payload regions read or written in order.
	// A batch operation requires their combined length to be nonzero.
	Buffers [][]byte
	// OOB contains Linux-compatible packet-info, hop-limit, traffic-class,
	// IPv6 flow-label, or asynchronous-error ancillary data.
	OOB []byte
	// Addr specifies the destination for an unconnected write and receives the
	// source address after a successful read. It must be nil for a connected
	// write.
	Addr net.Addr
	// N is the number of payload bytes read or written through Buffers.
	N int
	// NN is the number of ancillary bytes read or written through OOB.
	NN int
	// Flags contains Linux-compatible MSG_TRUNC, MSG_CTRUNC, and MSG_ERRQUEUE
	// results after a successful read.
	Flags int
}

SocketMessage represents one scatter/gather datagram or IP protocol message. Its layout and field meanings match golang.org/x/net/ipv4.Message and golang.org/x/net/ipv6.Message without requiring either package.

type SocketOption

type SocketOption interface {
	// contains filtered or unexported methods
}

SocketOption is one strongly typed socket policy. Concrete option types are private; use SocketOptions to construct values understood by this package. Each constructor documents the creation operations that accept its value. Using an explicit setting outside that scope reports syscall.ENOPROTOOPT before an endpoint is created. Unset markers are accepted by every creation operation and have no effect where their corresponding setting is inapplicable.

type SocketOptionFactory

type SocketOptionFactory uint8

SocketOptionFactory constructs socket options used when creating sockets. Its value carries no state; use SocketOptions rather than constructing one.

const SocketOptions SocketOptionFactory = 0

SocketOptions constructs creation-time socket policies. See the methods on SocketOptionFactory for each constructor's detailed contract. The constructors have the following operation scopes:

  • ReadBuffer, TrafficClass, and FlowLabel are valid for every creation method on ListenConfig and Dialer, TCPForwarderRequest.Accept, and UDPForwarderRequest.Accept or UDPForwarderRequest.Listen.
  • WriteBuffer, KeepAlive, KeepAliveConfig, NoDelay, IdleTimeout, UserTimeout, CongestionControl, CongestionControlFactory, and MaximumPacingRate are valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.
  • AcceptQueue and SYNBacklog are valid only for ListenConfig.ListenTCP.
  • ReceiveErrors, PathMTUDiscovery, HopLimit, Broadcast, MulticastHopLimit, and MulticastLoopback are valid for the UDP and IP creation methods on ListenConfig and Dialer, UDPForwarderRequest.Accept, and UDPForwarderRequest.Listen.
  • ReuseAddress and ReusePort are valid only for ListenConfig.ListenTCP and ListenConfig.ListenUDP.
  • IPHeaderIncludedOnWrite, IPHeaderIncludedOnRead, ICMPv4Filter, ICMPv6Filter, and IPv6Checksum are valid only for ListenConfig.ListenIP and Dialer.DialIP. The latter three are also validated against the resolved address family and IP protocol.

Every Unset constructor is valid for every socket creation operation. It removes an applicable earlier override of the same kind and otherwise has no effect. Using an explicit setting outside its scope reports syscall.ENOPROTOOPT before an endpoint is created.

func (SocketOptionFactory) AcceptQueue

func (SocketOptionFactory) AcceptQueue(capacity int) SocketOption

AcceptQueue sets the number of completed TCP handshakes that may wait for Accept. Zero creates an unbuffered accept handoff; the value cannot be negative. The policy belongs to the listener and is not inherited by a connection. It is valid only for ListenConfig.ListenTCP.

func (SocketOptionFactory) Broadcast

func (SocketOptionFactory) Broadcast(enabled bool) SocketOption

Broadcast controls the SO_BROADCAST-equivalent output permission inherited by newly created UDP and IP sockets. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) CongestionControl

func (SocketOptionFactory) CongestionControl(algorithm string) SocketOption

CongestionControl selects one registered algorithm by name for newly created TCP connections. It overrides CongestionControlFactory when it appears later in the same option list. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) CongestionControlFactory

func (SocketOptionFactory) CongestionControlFactory(factory *CongestionControlFactory) SocketOption

CongestionControlFactory selects an immutable local factory for newly created TCP connections. It overrides CongestionControl when it appears later in the same option list. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) FlowLabel

func (SocketOptionFactory) FlowLabel(label uint32) SocketOption

FlowLabel fixes the IPv6 Flow Label inherited by a newly created TCP, UDP, or IP socket. Zero explicitly disables automatic labeling. An IPv4-only endpoint reports EAFNOSUPPORT when the socket is created. Label must fit in 20 bits.

It is valid for every creation method on ListenConfig and Dialer, for TCPForwarderRequest.Accept, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) HopLimit

func (SocketOptionFactory) HopLimit(hopLimit int) SocketOption

HopLimit sets the default unicast IPv4 TTL or IPv6 Hop Limit inherited by newly created UDP and IP sockets. Value must be in [0, 255]. Zero is valid only for an IPv6-only endpoint; IPv4 and dual-stack creation report EINVAL. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) ICMPv4Filter

func (SocketOptionFactory) ICMPv4Filter(filter ICMPv4Filter) SocketOption

ICMPv4Filter installs a receive-type filter on a newly created IPv4 ICMP protocol socket. A generic dual-stack ip:icmp socket applies it only to its IPv4 branch. Other protocols report syscall.ENOPROTOOPT and IPv6-only sockets report syscall.EAFNOSUPPORT before the endpoint is created.

func (SocketOptionFactory) ICMPv6Filter

func (SocketOptionFactory) ICMPv6Filter(filter ICMPv6Filter) SocketOption

ICMPv6Filter installs a receive-type filter on a newly created ICMPv6 protocol socket. A generic dual-stack ip:ipv6-icmp socket applies it only to its IPv6 branch. Other protocols report syscall.ENOPROTOOPT and IPv4-only sockets report syscall.EAFNOSUPPORT before the endpoint is created.

func (SocketOptionFactory) IPHeaderIncludedOnRead

func (SocketOptionFactory) IPHeaderIncludedOnRead(enabled bool) SocketOption

IPHeaderIncludedOnRead controls whether IPConn reads return the complete, reassembled IP packet instead of only its protocol payload. It is a creation-time option because changing the interpretation of queued messages would make concurrent reads ambiguous. It is valid only for ListenConfig.ListenIP and Dialer.DialIP.

func (SocketOptionFactory) IPHeaderIncludedOnWrite

func (SocketOptionFactory) IPHeaderIncludedOnWrite(enabled bool) SocketOption

IPHeaderIncludedOnWrite controls whether IPConn writes contain a complete IPv4 or IPv6 packet instead of a protocol payload. It corresponds to IP_HDRINCL and IPV6_HDRINCL. Use IPConn.SetIPHeaderIncludedOnWrite to change the representation of an existing socket. It is valid only for ListenConfig.ListenIP and Dialer.DialIP.

func (SocketOptionFactory) IPv6Checksum

func (SocketOptionFactory) IPv6Checksum(enabled bool, offset int) SocketOption

IPv6Checksum controls RFC 3542 IPV6_CHECKSUM processing for a newly created non-ICMPv6 protocol socket. When enabled, offset is the even, non-negative byte offset of the 16-bit checksum field in the upper-layer payload. When disabled, offset is ignored. IPv4-only sockets report syscall.EAFNOSUPPORT; ICMPv6 sockets report syscall.EINVAL because their checksum at offset 2 is mandatory and cannot be configured.

func (SocketOptionFactory) IdleTimeout

func (SocketOptionFactory) IdleTimeout(timeout time.Duration) SocketOption

IdleTimeout closes newly created TCP connections after this duration without an acceptable inbound segment. Zero explicitly disables the policy. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) KeepAlive

func (SocketOptionFactory) KeepAlive(enabled bool) SocketOption

KeepAlive controls keepalive probing on newly created TCP connections. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) KeepAliveConfig

func (SocketOptionFactory) KeepAliveConfig(config KeepAliveConfig) SocketOption

KeepAliveConfig sets the keepalive idle interval, probe interval, and probe count inherited by newly created TCP connections. Every field must be positive. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) MaximumPacingRate

func (SocketOptionFactory) MaximumPacingRate(bytesPerSecond uint64) SocketOption

MaximumPacingRate caps newly created TCP connections to bytesPerSecond. Zero explicitly removes the cap. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) MulticastHopLimit

func (SocketOptionFactory) MulticastHopLimit(hopLimit int) SocketOption

MulticastHopLimit sets the IPv4 multicast TTL or IPv6 multicast Hop Limit inherited by newly created UDP and IP sockets. Value must be in [0, 255]; zero confines output to this host. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) MulticastLoopback

func (SocketOptionFactory) MulticastLoopback(enabled bool) SocketOption

MulticastLoopback controls delivery of transmitted multicast packets to matching local memberships for newly created UDP and IP sockets. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) NoDelay

func (SocketOptionFactory) NoDelay(enabled bool) SocketOption

NoDelay controls Nagle coalescing on newly created TCP connections. True is the package and net.TCPConn-compatible default. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) PathMTUDiscovery

func (SocketOptionFactory) PathMTUDiscovery(mode PathMTUDiscovery) SocketOption

PathMTUDiscovery sets the Linux-compatible IP_MTU_DISCOVER policy inherited by newly created UDP and IP sockets. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) ReadBuffer

func (SocketOptionFactory) ReadBuffer(bytes int) SocketOption

ReadBuffer fixes the receive-buffer capacity of a newly created TCP, UDP, or IP socket. On TCP it disables receive auto-tuning for that connection. The value must be positive; UDP and IP raise values below one message's metadata cost to that protocol's minimum usable capacity.

It is valid for every creation method on ListenConfig and Dialer, for TCPForwarderRequest.Accept, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) ReceiveErrors

func (SocketOptionFactory) ReceiveErrors(enabled bool) SocketOption

ReceiveErrors controls whether newly created UDP and IP sockets retain reportable asynchronous errors for ReadError. Ordinary reads and UDP writes still consume pending socket errors; IP writes do so only for header-included packets. Connected sockets report hard errors before queued payloads, and enabled sockets also report eligible soft errors. It also makes writes return ENOBUFS when immediate admission of unicast output or the external-link copy of multicast or broadcast output fails. It does not report packets displaced after admission. Receive-side non-unicast loopback copies remain best effort. It is valid for the UDP and IP creation methods on ListenConfig and Dialer, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) ReuseAddress

func (SocketOptionFactory) ReuseAddress(enabled bool) SocketOption

ReuseAddress controls Linux SO_REUSEADDR-style address reuse during bind. TCP listeners enable it by default, matching Go's standard listener setup; an explicit false value disables that behavior. UDP listeners default to an exclusive binding and require every overlapping endpoint to opt in. It is valid only for ListenConfig.ListenTCP and ListenConfig.ListenUDP.

func (SocketOptionFactory) ReusePort

func (SocketOptionFactory) ReusePort(enabled bool) SocketOption

ReusePort controls Linux SO_REUSEPORT-style flow distribution for TCP and UDP listeners. Every endpoint in an overlapping group must enable it. It is valid only for ListenConfig.ListenTCP and ListenConfig.ListenUDP.

func (SocketOptionFactory) SYNBacklog

func (SocketOptionFactory) SYNBacklog(capacity int) SocketOption

SYNBacklog sets the number of stateful TCP handshakes owned by a listener before it falls back to SYN cookies. Zero selects cookies immediately; the value cannot be negative. It is valid only for ListenConfig.ListenTCP.

func (SocketOptionFactory) TrafficClass

func (SocketOptionFactory) TrafficClass(value int) SocketOption

TrafficClass sets the IPv4 TOS or IPv6 Traffic Class byte inherited by a newly created TCP, UDP, or IP socket. TCP masks the two ECN bits because its transport state controls them independently. Value must be in [0, 255].

It is valid for every creation method on ListenConfig and Dialer, for TCPForwarderRequest.Accept, and for UDPForwarderRequest.Accept and UDPForwarderRequest.Listen.

func (SocketOptionFactory) UnsetAcceptQueue

func (SocketOptionFactory) UnsetAcceptQueue() SocketOption

UnsetAcceptQueue restores the current Stack completed-handshake limit. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetBroadcast

func (SocketOptionFactory) UnsetBroadcast() SocketOption

UnsetBroadcast restores the current Stack broadcast-output policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetCongestionControl

func (SocketOptionFactory) UnsetCongestionControl() SocketOption

UnsetCongestionControl restores the current Stack congestion-control policy, overriding either congestion-control option form used earlier. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetFlowLabel

func (SocketOptionFactory) UnsetFlowLabel() SocketOption

UnsetFlowLabel restores inheritance, including automatic flow-label selection when the Stack default is zero. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetHopLimit

func (SocketOptionFactory) UnsetHopLimit() SocketOption

UnsetHopLimit restores the current Stack unicast hop-limit policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetICMPv4Filter

func (SocketOptionFactory) UnsetICMPv4Filter() SocketOption

UnsetICMPv4Filter restores the all-accepting default, overriding an earlier ICMPv4Filter option in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetICMPv6Filter

func (SocketOptionFactory) UnsetICMPv6Filter() SocketOption

UnsetICMPv6Filter restores the all-accepting default, overriding an earlier ICMPv6Filter option in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetIPHeaderIncludedOnRead

func (SocketOptionFactory) UnsetIPHeaderIncludedOnRead() SocketOption

UnsetIPHeaderIncludedOnRead restores the current Stack IP read default, overriding earlier IPHeaderIncludedOnRead options in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetIPHeaderIncludedOnWrite

func (SocketOptionFactory) UnsetIPHeaderIncludedOnWrite() SocketOption

UnsetIPHeaderIncludedOnWrite restores the current Stack IP write default, overriding earlier IPHeaderIncludedOnWrite options in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetIPv6Checksum

func (SocketOptionFactory) UnsetIPv6Checksum() SocketOption

UnsetIPv6Checksum restores the protocol default, overriding an earlier IPv6Checksum option in the same list. ICMPv6 restores mandatory processing at offset 2; other protocols restore disabled processing. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetIdleTimeout

func (SocketOptionFactory) UnsetIdleTimeout() SocketOption

UnsetIdleTimeout restores the current Stack receive-idle policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetKeepAlive

func (SocketOptionFactory) UnsetKeepAlive() SocketOption

UnsetKeepAlive restores the current Stack TCP keepalive default. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetKeepAliveConfig

func (SocketOptionFactory) UnsetKeepAliveConfig() SocketOption

UnsetKeepAliveConfig restores the current Stack keepalive timing policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetMaximumPacingRate

func (SocketOptionFactory) UnsetMaximumPacingRate() SocketOption

UnsetMaximumPacingRate restores the current Stack pacing-rate policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetMulticastHopLimit

func (SocketOptionFactory) UnsetMulticastHopLimit() SocketOption

UnsetMulticastHopLimit restores the current Stack multicast hop limit. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetMulticastLoopback

func (SocketOptionFactory) UnsetMulticastLoopback() SocketOption

UnsetMulticastLoopback restores the current Stack multicast-loopback policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetNoDelay

func (SocketOptionFactory) UnsetNoDelay() SocketOption

UnsetNoDelay restores the current Stack TCP Nagle policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetPathMTUDiscovery

func (SocketOptionFactory) UnsetPathMTUDiscovery() SocketOption

UnsetPathMTUDiscovery restores the current Stack PMTU-discovery policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetReadBuffer

func (SocketOptionFactory) UnsetReadBuffer() SocketOption

UnsetReadBuffer restores inheritance from the current Stack configuration, overriding an earlier ReadBuffer option in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetReceiveErrors

func (SocketOptionFactory) UnsetReceiveErrors() SocketOption

UnsetReceiveErrors restores the current Stack ReceiveErrors policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetReuseAddress

func (SocketOptionFactory) UnsetReuseAddress() SocketOption

UnsetReuseAddress restores the operation-specific SO_REUSEADDR default, overriding earlier ReuseAddress options in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetReusePort

func (SocketOptionFactory) UnsetReusePort() SocketOption

UnsetReusePort restores the default disabled SO_REUSEPORT policy, overriding earlier ReusePort options in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetSYNBacklog

func (SocketOptionFactory) UnsetSYNBacklog() SocketOption

UnsetSYNBacklog restores the current Stack stateful-handshake limit. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetTrafficClass

func (SocketOptionFactory) UnsetTrafficClass() SocketOption

UnsetTrafficClass restores inheritance from the current Stack configuration, overriding an earlier TrafficClass option in the same list. The unset marker is valid for every socket creation operation.

func (SocketOptionFactory) UnsetUserTimeout

func (SocketOptionFactory) UnsetUserTimeout() SocketOption

UnsetUserTimeout restores the current Stack TCP user-timeout policy. It is valid for every socket creation operation.

func (SocketOptionFactory) UnsetWriteBuffer

func (SocketOptionFactory) UnsetWriteBuffer() SocketOption

UnsetWriteBuffer restores TCP send-buffer inheritance and auto-tuning. It is valid for every socket creation operation.

func (SocketOptionFactory) UserTimeout

func (SocketOptionFactory) UserTimeout(timeout time.Duration) SocketOption

UserTimeout bounds how long data on a newly created TCP connection may remain unacknowledged or unsent behind a zero window. Zero explicitly disables this custom bound. It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

func (SocketOptionFactory) WriteBuffer

func (SocketOptionFactory) WriteBuffer(bytes int) SocketOption

WriteBuffer fixes the send-buffer capacity of a newly created TCP connection and disables send auto-tuning for that connection. It is not a UDP or IP option because those protocols make one immediate attempt to admit output to a bounded stack queue and retain no per-socket send buffer.

It is valid for ListenConfig.ListenTCP, Dialer.DialTCP, and TCPForwarderRequest.Accept.

type Stack

type Stack struct {
	// contains filtered or unexported fields
}

Stack converts raw IPv4/IPv6 packets to application TCP, UDP, and IP protocol sockets.

func New

func New(config Config) (*Stack, error)

New constructs an inactive-socket stack.

func (*Stack) BatchSize

func (s *Stack) BatchSize() int

BatchSize reports the maximum number of packets returned by one Read call. Write accepts larger batches so Stack remains compatible with composite packet devices whose network side has a larger batch size.

func (*Stack) Close

func (s *Stack) Close() error

Close cancels packet-device I/O, rejects later calls, and starts orderly socket and background-worker shutdown. It does not wait for socket actors or user forwarder handlers to return.

func (*Stack) ConfirmPathMTU

func (s *Stack) ConfirmPathMTU(destination netip.Addr, mtu int) error

ConfirmPathMTU records packetization-layer acknowledgement of an unfragmented probe. Connectionless protocols must call this only after their own acknowledgement semantics prove delivery; queueing a packet is not confirmation.

func (*Stack) DialIP

func (s *Stack) DialIP(ctx context.Context, network string, source, remote netip.Addr) (net.Conn, error)

DialIP creates a connected IPv4 or IPv6 protocol socket. Network must be an IP network with a numeric or well-known protocol, such as ip6:ipv6-icmp or ip:99. An invalid or unspecified source selects a managed address using the route table.

func (*Stack) DialTCP

func (s *Stack) DialTCP(ctx context.Context, network string, source, remote netip.AddrPort) (net.Conn, error)

DialTCP establishes an active IPv4 or IPv6 TCP connection. Network must be tcp, tcp4, or tcp6. A zero source selects both address and port automatically; an unspecified source address selects only the address.

func (*Stack) DialUDP

func (s *Stack) DialUDP(ctx context.Context, network string, source, remote netip.AddrPort) (net.Conn, error)

DialUDP creates a connected UDP socket for one IPv4 or IPv6 remote endpoint. Network must be udp, udp4, or udp6. A zero source selects both address and port automatically; an unspecified source address selects only the address.

func (*Stack) ListenIP

func (s *Stack) ListenIP(ctx context.Context, network string, local netip.Addr) (net.PacketConn, error)

ListenIP creates an unconnected IPv4 or IPv6 protocol socket. Network must be an IP network with a numeric or well-known protocol, such as ip4:icmp or ip:99. An empty Local selects the network's wildcard address; a generic ip wildcard is dual-stack when both address families are configured. The returned net.PacketConn has dynamic type *IPConn.

func (*Stack) ListenMulticastUDP

func (s *Stack) ListenMulticastUDP(ctx context.Context, network string, group netip.AddrPort) (*UDPConn, error)

ListenMulticastUDP creates a reusable UDP socket, disables multicast loopback like net.ListenMulticastUDP, and joins group. Mipstack has one embedding interface, so no interface selector is required. Source-specific groups return EINVAL; use ListenConfig with SocketOptions.ReuseAddress and JoinSourceSpecificGroup.

func (*Stack) ListenTCP

func (s *Stack) ListenTCP(ctx context.Context, network string, local netip.AddrPort) (net.Listener, error)

ListenTCP creates a passive TCP endpoint. Network must be tcp, tcp4, or tcp6. A wildcard with tcp uses one dual-stack endpoint when both families are configured. Port zero selects an automatic port. The returned net.Listener has dynamic type *TCPListener.

func (*Stack) ListenUDP

func (s *Stack) ListenUDP(ctx context.Context, network string, local netip.AddrPort) (net.PacketConn, error)

ListenUDP binds an unconnected UDP packet socket. Network must be udp, udp4, or udp6. A wildcard with udp uses one dual-stack endpoint when both families are configured. Port zero selects an automatic port. The returned net.PacketConn has dynamic type *UDPConn.

func (*Stack) LocalAddresses

func (s *Stack) LocalAddresses() []netip.Addr

LocalAddresses returns an independent snapshot of all configured local addresses in configuration order.

func (*Stack) MTU

func (s *Stack) MTU() (int, error)

MTU returns the current link MTU.

func (*Stack) Name

func (s *Stack) Name() (string, error)

Name returns the stable packet-device implementation name.

func (*Stack) PathMTU

func (s *Stack) PathMTU(destination netip.Addr) (int, error)

PathMTU returns the currently confirmed packet size for one routed unicast destination. The result includes the IP header.

func (*Stack) RXChecksumOffload

func (s *Stack) RXChecksumOffload() RXChecksumOffload

RXChecksumOffload returns an independent copy of the input link's checksum verification policy.

func (*Stack) Read

func (s *Stack) Read(buffers [][]byte, sizes []int, offset int) (int, error)

Read copies complete outbound IP packets into consecutive buffers beginning at offset. It blocks for the first packet, then drains only packets that are already ready, returning at most BatchSize packets. It writes lengths only to sizes[:n]; later elements are unchanged and must be ignored.

If a destination buffer is too short, Read discards that packet and returns io.ErrShortBuffer together with the number of earlier packets copied by the call. Other errors likewise return the successfully completed packet prefix. Close unblocks a waiting Read with os.ErrClosed.

Outbound capacity is finite. If the embedding device stops calling Read, TCP retains protocol work until capacity returns while its socket send-buffer and deadline rules remain in force. UDP and IP writes make one immediate admission attempt. Under overload, nonblocking datagram and control output may displace queued packets or be discarded while the scheduler preserves progress across flows. Failure to admit unicast output or an external-link non-unicast copy is successful by default and reports ENOBUFS when the socket's ReceiveErrors policy is enabled; that policy does not report later displacement. Receive-side non-unicast loopback copies remain independently best effort. A successful socket write therefore does not guarantee that every resulting packet will be returned by Read. Resuming Read releases capacity for pending TCP work.

Read may run concurrently with Write and with other Read calls. Each queued packet is assigned to at most one call, but concurrent calls have no relative completion order. Stack does not access the destination buffers after Read returns, so the caller may reuse them immediately.

func (*Stack) RouteFor

func (s *Stack) RouteFor(destination netip.Addr) (Route, error)

RouteFor returns the selected route for one unicast destination.

func (*Stack) SetRXChecksumOffload

func (s *Stack) SetRXChecksumOffload(offload RXChecksumOffload)

SetRXChecksumOffload replaces the input link's checksum verification policy. Configure it before delivering packets; changes do not establish a processing boundary for input already in progress. It does not affect socket defaults.

func (*Stack) Start

func (s *Stack) Start() error

Start activates packet and socket I/O and starts background maintenance. Repeated calls do not start additional workers.

func (*Stack) Stats

func (s *Stack) Stats() StackStats

Stats returns a consistent-enough lock-free snapshot of stack counters. Concurrent activity may become visible across adjacent fields at slightly different instants.

func (*Stack) UpdateConfig

func (s *Stack) UpdateConfig(config Config) error

UpdateConfig validates and atomically replaces the stack configuration. New sockets inherit the replacement defaults. Sockets bound to removed addresses or destinations without a remaining route are closed. Other TCP connections immediately apply a changed default congestion controller and reclamp their MSS; existing sockets retain the remaining inherited policies.

func (*Stack) Write

func (s *Stack) Write(buffers [][]byte, offset int) (int, error)

Write consumes complete inbound IP packets from buffers beginning at offset, in slice order. It accepts any number of buffers and is not limited by BatchSize. Invalid, unrelated, and unsupported packets are accounted as drops but still count as successfully consumed and do not produce an error.

An error after one or more buffers returns the successfully completed packet prefix. Close causes pending or subsequent work to return os.ErrClosed. Write may run concurrently with Read and with other Write calls; concurrent calls have no relative processing order. Stack retains no reference to the buffers after Write returns, so the caller may reuse them immediately.

type StackStats

type StackStats struct {
	// InboundPackets counts complete packets presented to the stack.
	InboundPackets uint64
	// InboundDroppedPackets counts invalid packets and bounded-queue drops.
	InboundDroppedPackets uint64
	// InvalidIPPackets counts packets rejected by IP parsing or reassembly.
	InvalidIPPackets uint64
	// UnacceptedIPPackets counts valid packets whose source or destination is
	// not admissible for this endpoint stack.
	UnacceptedIPPackets uint64
	// NonlocalDestinationPackets is the unaccepted subset addressed elsewhere.
	NonlocalDestinationPackets uint64
	// PromiscuousInboundPackets counts valid packets admitted for a destination
	// not present in LocalAddresses.
	PromiscuousInboundPackets uint64
	// InvalidSourcePackets is the unaccepted subset with a prohibited source.
	InvalidSourcePackets uint64
	// OutboundPackets counts complete packets accepted by the device queue,
	// including packets later displaced under overload.
	OutboundPackets uint64
	// OutboundQueueDrops counts individual packets rejected by or displaced from
	// the bounded queue consumed by Read. It excludes shutdown cleanup and loss
	// after Read returns a packet.
	OutboundQueueDrops uint64
	// LoopbackPackets counts locally routed packets that bypassed the link.
	LoopbackPackets uint64
	// LoopbackQueueDrops counts individual packets rejected by the bounded local
	// delivery queue. An all-or-none fragmented sequence counts each rejected
	// fragment; shutdown cleanup is excluded.
	LoopbackQueueDrops uint64
	// ActiveTCPConnections includes handshakes, established flows, and
	// TIME_WAIT actors.
	ActiveTCPConnections uint64
	// ActiveTCPListeners is the current number of passive TCP endpoints.
	ActiveTCPListeners uint64
	// ActiveUDPSockets is the current number of open packet sockets.
	ActiveUDPSockets uint64
	// ActiveIPSockets is the current number of open protocol sockets.
	ActiveIPSockets uint64
	// TCPRetransmissions counts all SYN, data, FIN, SACK, RACK, and tail-probe
	// retransmissions.
	TCPRetransmissions uint64
	// TCPInboundQueueDrops counts validated segments rejected by a connection's
	// byte-bounded actor queue.
	TCPInboundQueueDrops uint64
	// TCPInvalidSegments counts malformed headers and checksum failures.
	TCPInvalidSegments uint64
	// TCPSACKRetransmissions counts retransmissions selected by the SACK
	// scoreboard, including its RACK-confirmed subset.
	TCPSACKRetransmissions uint64
	// TCPRACKRetransmissions counts the time-based subset of SACK recovery.
	TCPRACKRetransmissions uint64
	// TCPTailLossProbes counts probes sent before the ordinary RTO.
	TCPTailLossProbes uint64
	// TCPSpuriousRecoveryUndos counts Eifel, DSACK, or F-RTO evidence that
	// safely restored congestion state after an unnecessary retransmission.
	TCPSpuriousRecoveryUndos uint64
	// TCPZeroWindowProbes counts persist probes sent while the peer advertises
	// a closed receive window.
	TCPZeroWindowProbes uint64
	// TCPKeepAliveProbes counts probes sent after configured receive inactivity.
	TCPKeepAliveProbes uint64
	// TCPSYNCookiesSent counts stateless SYN-ACKs emitted under listener or
	// stack connection pressure.
	TCPSYNCookiesSent uint64
	// TCPSYNCookiesAccepted counts final ACKs that authenticated a recent SYN
	// cookie and entered a listener backlog.
	TCPSYNCookiesAccepted uint64
	// TCPSYNCookiesRejected counts candidate final ACKs that failed cookie
	// authentication while cookie validation was active.
	TCPSYNCookiesRejected uint64
	// TCPHandshakeTimeouts counts passive stateful handshakes that exhausted
	// their SYN-ACK retry budget.
	TCPHandshakeTimeouts uint64
	// TCPAcceptQueueDrops counts completed handshakes aborted because their
	// listener's accept queue was full.
	TCPAcceptQueueDrops uint64
	// PathMTUUpdates counts accepted destination PMTU reductions.
	PathMTUUpdates uint64
	// PathMTUProbes counts TCP packets sent above the confirmed effective MTU.
	PathMTUProbes uint64
	// PathMTUProbeSuccesses counts acknowledged upward TCP probes.
	PathMTUProbeSuccesses uint64
	// PathMTUProbeFailures counts isolated upward probes rejected by SACK
	// evidence without treating them as congestion loss.
	PathMTUProbeFailures uint64
	// PathMTUBlackHoleReductions counts PMTU reductions inferred from repeated
	// TCP timeouts rather than ICMP.
	PathMTUBlackHoleReductions uint64
	// FragmentEvictions counts incomplete datagrams removed for capacity.
	FragmentEvictions uint64
	// FragmentTimeouts counts incomplete datagrams removed for age.
	FragmentTimeouts uint64
	// RateLimitedControlResponses counts suppressed TCP RST and challenge ACK,
	// ICMP unreachable, parameter-problem, and ICMP echo replies.
	RateLimitedControlResponses uint64
}

StackStats is a point-in-time snapshot of stack activity. Counters are monotonic except ActiveTCPConnections, ActiveTCPListeners, ActiveUDPSockets, and ActiveIPSockets.

type TCPConn

type TCPConn struct {
	// contains filtered or unexported fields
}

TCPConn is an active userspace TCP connection.

func (*TCPConn) Close

func (c *TCPConn) Close() error

Close releases application access and applies the SetLinger policy. The default queues FIN after accepted writes and finishes protocol processing in the background.

func (*TCPConn) CloseRead

func (c *TCPConn) CloseRead() error

CloseRead closes the application receive direction without resetting TCP.

func (*TCPConn) CloseWrite

func (c *TCPConn) CloseWrite() error

CloseWrite queues FIN after all bytes already accepted by Write.

func (*TCPConn) Info

func (c *TCPConn) Info() TCPConnInfo

Info returns a consistent diagnostic snapshot. The connection actor supplies live protocol state; after termination, the final snapshot remains available for post-mortem inspection.

func (*TCPConn) LocalAddr

func (c *TCPConn) LocalAddr() net.Addr

LocalAddr returns the managed local TCP endpoint.

func (*TCPConn) MultipathTCP

func (c *TCPConn) MultipathTCP() (bool, error)

MultipathTCP reports whether this connection uses MPTCP. Mipstack currently implements ordinary TCP only, so the result is always false.

func (*TCPConn) Read

func (c *TCPConn) Read(buffer []byte) (int, error)

Read returns contiguous application bytes or the receive terminal state.

func (*TCPConn) ReadFrom

func (c *TCPConn) ReadFrom(reader io.Reader) (int64, error)

ReadFrom copies a stream into c while preserving write ordering with concurrent calls to Write. It implements io.ReaderFrom without recursively entering io.Copy's ReaderFrom fast path.

func (*TCPConn) ReadWithBuffer

func (c *TCPConn) ReadWithBuffer(getBuffer func(sizeHint int) []byte) (int, error)

ReadWithBuffer reads contiguous application bytes like Read, obtaining the destination buffer lazily from getBuffer.

If application data is available, getBuffer is called once after this method observes it. It is not called when the operation returns before data is available because of an error or deadline. The callback receives the currently queued application byte count as an advisory size hint, which may be stale when it returns. It runs without c's connection state lock and must return promptly; it must not call Read, ReadWithBuffer, or WriteTo on c. The caller owns the returned slice, and c does not retain it. The slice length limits the number of bytes read. A nil callback returns EINVAL. An empty returned slice reports io.ErrShortBuffer without consuming receive data. If the callback is called, the caller must release the returned buffer even when this method returns an error.

This is an experimental API and is not covered by the package's stability guarantees.

func (*TCPConn) RemoteAddr

func (c *TCPConn) RemoteAddr() net.Addr

RemoteAddr returns the connected remote TCP endpoint.

func (*TCPConn) SetCongestionControl

func (c *TCPConn) SetCongestionControl(algorithm string) error

SetCongestionControl selects a registered algorithm by name for this connection and prevents later stack-default updates from overriding the explicit choice.

func (*TCPConn) SetCongestionControlFactory

func (c *TCPConn) SetCongestionControlFactory(factory *CongestionControlFactory) error

SetCongestionControlFactory changes this connection to an immutable local factory and prevents later stack-default updates from overriding the explicit choice. The factory creates a new connection-private controller on the actor goroutine; it may be shared with other connections safely.

func (*TCPConn) SetDeadline

func (c *TCPConn) SetDeadline(deadline time.Time) error

SetDeadline updates both application deadlines.

func (*TCPConn) SetIdleTimeout

func (c *TCPConn) SetIdleTimeout(timeout time.Duration) error

SetIdleTimeout closes the connection when no acceptable segment arrives for timeout. Zero disables the timeout.

func (*TCPConn) SetKeepAlive

func (c *TCPConn) SetKeepAlive(enabled bool) error

SetKeepAlive enables or disables TCP keepalive probes.

func (*TCPConn) SetKeepAliveConfig

func (c *TCPConn) SetKeepAliveConfig(config KeepAliveConfig) error

SetKeepAliveConfig replaces keepalive timing and probe count.

func (*TCPConn) SetKeepAlivePeriod

func (c *TCPConn) SetKeepAlivePeriod(period time.Duration) error

SetKeepAlivePeriod sets both the idle delay and probe interval. Use SetKeepAliveConfig when different values or a custom probe count are needed.

func (*TCPConn) SetLinger

func (c *TCPConn) SetLinger(seconds int) error

SetLinger controls how Close handles data waiting to be sent or acknowledged. A negative value completes in the background, zero performs an abortive close, and a positive value waits up to that many seconds before aborting the remaining transmission.

func (*TCPConn) SetMaximumPacingRate

func (c *TCPConn) SetMaximumPacingRate(bytesPerSecond uint64) error

SetMaximumPacingRate caps this connection's paced-data rate in bytes per second. Zero removes the limit. Initial and control bursts mean it is not a strict byte-rate shaper. The selected congestion controller still maintains its unconstrained path model so removing a limit takes effect without resetting congestion state.

func (*TCPConn) SetNoDelay

func (c *TCPConn) SetNoDelay(noDelay bool) error

SetNoDelay controls Nagle coalescing. The default is true, matching net.TCPConn.

func (*TCPConn) SetQuickACK

func (c *TCPConn) SetQuickACK(enabled bool) error

SetQuickACK requests Linux TCP_QUICKACK-style receive behavior. Enabling it replenishes a bounded immediate-ACK budget, leaves response-piggybacking mode, and flushes a pending acknowledgement; disabling it enters response-piggybacking mode. Protocol events may still require prompt feedback or change the mode, so the request is not persistent. A successful call queues the actor-owned policy change; it does not wait for an acknowledgement to reach the link.

func (*TCPConn) SetReadBuffer

func (c *TCPConn) SetReadBuffer(bytes int) error

SetReadBuffer changes the bounded application receive capacity.

func (*TCPConn) SetReadDeadline

func (c *TCPConn) SetReadDeadline(deadline time.Time) error

SetReadDeadline updates the next Read deadline.

func (*TCPConn) SetTrafficClass

func (c *TCPConn) SetTrafficClass(value int) error

SetTrafficClass sets IPv4 TOS or IPv6 Traffic Class DSCP bits. TCP owns and replaces the two ECN bits on each packet.

func (*TCPConn) SetUserTimeout

func (c *TCPConn) SetUserTimeout(timeout time.Duration) error

SetUserTimeout bounds how long transmitted data may remain unacknowledged, or buffered data may remain unsent behind a zero window. Zero disables the custom bound while retaining the normal TCP retry limits. Like Linux TCP_USER_TIMEOUT, this is a local policy and does not negotiate the RFC 5482 UTO option.

func (*TCPConn) SetWriteBuffer

func (c *TCPConn) SetWriteBuffer(bytes int) error

SetWriteBuffer changes the bounded application send capacity.

func (*TCPConn) SetWriteDeadline

func (c *TCPConn) SetWriteDeadline(deadline time.Time) error

SetWriteDeadline updates the next Write deadline.

func (*TCPConn) Write

func (c *TCPConn) Write(payload []byte) (int, error)

Write copies payload into the bounded TCP send buffer. It waits only for buffer space, not peer acknowledgement, matching standard net.Conn semantics. Bytes reported as written remain queued after a later timeout.

func (*TCPConn) WriteTo

func (c *TCPConn) WriteTo(writer io.Writer) (int64, error)

WriteTo copies the receive stream into writer while preserving read ordering with concurrent calls to Read. It implements io.WriterTo.

type TCPConnInfo

type TCPConnInfo struct {
	// LocalAddress is the local TCP endpoint.
	LocalAddress netip.AddrPort
	// RemoteAddress is the peer TCP endpoint.
	RemoteAddress netip.AddrPort
	// State is the current RFC 9293 connection state.
	State TCPState
	// CongestionControl is the selected controller's diagnostic name.
	CongestionControl string
	// RTT is the smoothed round-trip time.
	RTT time.Duration
	// MinimumRTT is the minimum recent round-trip time.
	MinimumRTT time.Duration
	// RTTVariation is the smoothed round-trip-time variation.
	RTTVariation time.Duration
	// RetransmissionTimeout is the current RFC 6298 RTO.
	RetransmissionTimeout time.Duration
	// CongestionWindow is the current congestion window in bytes.
	CongestionWindow uint32
	// SlowStartThreshold is the current slow-start threshold in bytes.
	SlowStartThreshold uint32
	// BytesInFlight is the amount of transmitted but unacknowledged data.
	BytesInFlight uint32
	// DeliveryRate is the most recent delivery-rate estimate in bytes per second.
	DeliveryRate uint64
	// PacingRate is the controller's current pacing rate in bytes per second.
	PacingRate uint64
	// MaximumPacingRate is the configured pacing-rate ceiling, or zero if unlimited.
	MaximumPacingRate uint64
	// CongestionState is the controller-specific diagnostic state name.
	CongestionState string
	// PeerWindow is the latest advertised peer receive window in bytes.
	PeerWindow uint32
	// ReceiveWindow is the currently advertised local receive window in bytes.
	ReceiveWindow uint32
	// MaximumSegmentSize is the effective outbound TCP payload ceiling.
	MaximumSegmentSize int
	// PathMTU is the effective complete-IP-packet path MTU.
	PathMTU int
	// SendBufferSize is the number of application bytes currently buffered for send.
	SendBufferSize int
	// SendBufferCapacity is the current send-buffer limit.
	SendBufferCapacity int
	// MaximumSendBuffer is the automatic send-buffer tuning ceiling.
	MaximumSendBuffer int
	// ReceiveBufferSize is the number of application bytes waiting to be read.
	ReceiveBufferSize int
	// ReceiveBufferCapacity is the current receive-buffer limit.
	ReceiveBufferCapacity int
	// MaximumReceiveBuffer is the automatic receive-buffer tuning ceiling.
	MaximumReceiveBuffer int
	// BytesSent counts original application bytes transmitted.
	BytesSent uint64
	// BytesAcknowledged counts application bytes cumulatively acknowledged.
	BytesAcknowledged uint64
	// BytesReceived counts in-order application bytes received.
	BytesReceived uint64
	// Retransmissions counts retransmitted TCP segments.
	Retransmissions uint64
	// InboundQueueDrops counts segments rejected by the bounded actor queue.
	InboundQueueDrops uint64
	// InboundQueueBytes is the packet memory retained by the actor queue.
	InboundQueueBytes int64
	// InboundQueuePeak is the lifetime peak retained actor-queue memory.
	InboundQueuePeak int64
	// InboundQueueCapacity is the actor queue's memory bound.
	InboundQueueCapacity int
	// FastRecovery reports whether loss or ECN fast recovery is active.
	FastRecovery bool
	// RetransmissionRecovery reports whether recovery was entered by an RTO.
	RetransmissionRecovery bool
	// HyStartCSS reports whether HyStart++ Conservative Slow Start is active.
	HyStartCSS bool
	// PathMTUDiscovery reports whether packetization-layer discovery is enabled.
	PathMTUDiscovery bool
	// PathMTUProbe is the complete packet size of an outstanding probe, or zero.
	PathMTUProbe int
	// WindowScaling reports whether RFC 7323 window scaling was negotiated.
	WindowScaling bool
	// PeerWindowScale is the peer's advertised receive-window shift.
	PeerWindowScale uint8
	// ReceiveWindowScale is the local advertised receive-window shift.
	ReceiveWindowScale uint8
	// SACK reports whether selective acknowledgements were negotiated.
	SACK bool
	// Timestamps reports whether RFC 7323 timestamps were negotiated.
	Timestamps bool
	// ECN reports whether explicit congestion notification was negotiated.
	ECN bool
	// KeepAlive reports whether keepalive probing is enabled.
	KeepAlive bool
	// KeepAliveConfig is the effective keepalive probe policy.
	KeepAliveConfig KeepAliveConfig
	// IdleTimeout is the configured bidirectional inactivity timeout.
	IdleTimeout time.Duration
	// UserTimeout is the configured TCP user timeout.
	UserTimeout time.Duration
	// NoDelay reports whether Nagle coalescing is disabled.
	NoDelay bool
	// TrafficClass is the IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass uint8
	// FlowLabel is the IPv6 Flow Label, or zero for IPv4.
	FlowLabel uint32
	// SpuriousRecoveryUndos counts Eifel, DSACK, or F-RTO recovery reversals.
	SpuriousRecoveryUndos uint64
	// PathMTUProbes counts packetization-layer probes sent.
	PathMTUProbes uint64
	// PathMTUProbeSuccesses counts acknowledged path-MTU probes.
	PathMTUProbeSuccesses uint64
	// PathMTUProbeFailures counts probes inferred lost.
	PathMTUProbeFailures uint64
	// ApplicationLimited reports whether delivery sampling is application-limited.
	ApplicationLimited bool
	// SchedulerLimited reports whether local scheduling currently limits delivery.
	SchedulerLimited bool
	// SchedulerLimitedEvents counts transitions into scheduler-limited delivery.
	SchedulerLimitedEvents uint64
	// LastError is the most recently recorded socket or asynchronous network error.
	LastError error
}

TCPConnInfo is a consistent point-in-time diagnostic snapshot of one TCP connection. Window, congestion, buffer, and path sizes are measured in bytes; PathMTU and MaximumSegmentSize include and exclude IP/TCP headers, respectively.

type TCPForwarder

type TCPForwarder struct {
	// contains filtered or unexported fields
}

TCPForwarder owns the single fallback TCP handler installed on a stack. It may be installed before or after Stack.Start.

func NewTCPForwarder

func NewTCPForwarder(stack *Stack, options TCPForwarderOptions, handler TCPForwarderHandler) (*TCPForwarder, error)

NewTCPForwarder installs a fallback handler for otherwise unhandled TCP connection attempts. Only one TCP forwarder may be active per stack. Promiscuous mode is not required for unhandled traffic addressed to LocalAddresses; Config.Promiscuous is required only for nonlocal destination addresses. Installing a forwarder does not start the stack.

func (*TCPForwarder) Close

func (f *TCPForwarder) Close() error

Close removes the TCP fallback handler and invalidates undecided requests. It does not wait for running handlers to return. An action already claimed by a handler may finish, and accepted connections remain open.

func (TCPForwarder) Done

func (f TCPForwarder) Done() <-chan struct{}

Done is closed when the forwarder is closed directly or by Stack.Close.

func (*TCPForwarder) Info

func (f *TCPForwarder) Info() ForwarderInfo

Info returns one TCP forwarder diagnostic snapshot.

type TCPForwarderHandler

type TCPForwarderHandler func(*TCPForwarderRequest)

TCPForwarderHandler decides the fate of one valid, otherwise unhandled SYN. MIPS starts a separate goroutine for each unique request, so handlers may run concurrently and must synchronize shared state. A handler may block, but an undecided request occupies the forwarder's MaxInFlight capacity for the entire block. The handler must call Accept, Drop, or Reject before returning; returning without an action drops the request. Accept must be called and allowed to return in the handler's own call; neither the request nor an in-progress action may be handed to another goroutine. After Accept returns, the resulting TCPConn is independent of the request: the handler may return immediately, retain the connection, or hand the connection to another goroutine.

type TCPForwarderOptions

type TCPForwarderOptions struct {
	// MaxInFlight bounds requests on which the handler has not yet selected an
	// action. Zero uses the configured TCP SYN backlog.
	MaxInFlight int
}

TCPForwarderOptions configures interception of otherwise unhandled TCP connection attempts. Zero fields retain stack defaults.

type TCPForwarderRequest

type TCPForwarderRequest struct {
	// contains filtered or unexported fields
}

TCPForwarderRequest is one valid initial SYN that did not match an ordinary TCP connection or listener. Exactly one terminal action is permitted during its handler call. A repeated or invalidated action reports ErrForwarderRequestCompleted.

func (*TCPForwarderRequest) Accept

func (r *TCPForwarderRequest) Accept(ctx context.Context, options ...SocketOption) (*TCPConn, error)

Accept creates a passive TCP endpoint and blocks until the handshake completes, ctx is canceled, or the stack closes. The accepted connection preserves the original destination in LocalAddr and the sender in RemoteAddr. The handler must wait for Accept to return before returning itself. Once Accept returns, the connection is independent of both the request and the forwarder: the handler may return immediately or hand the connection to another goroutine, and closing the forwarder does not close it. Options are validated before the request is claimed and are not retained; an invalid option leaves the request available for another action. After validation succeeds, Accept consumes the request even when endpoint creation or the handshake returns an error.

func (*TCPForwarderRequest) Done

func (r *TCPForwarderRequest) Done() <-chan struct{}

Done is closed when the handler returns or the request is invalidated by a configuration update or forwarder closure. It lets a blocking TCP handler abandon external work before attempting its terminal action.

func (*TCPForwarderRequest) Drop

func (r *TCPForwarderRequest) Drop() error

Drop consumes the TCP request without packet I/O. It may wait briefly for forwarder bookkeeping but does not wait for the network or output queue.

func (*TCPForwarderRequest) Flow

Flow returns the original inbound TCP four-tuple. The method may be called only during the handler, but the returned value is an independent copy that may be retained.

func (*TCPForwarderRequest) Reject

func (r *TCPForwarderRequest) Reject() error

Reject consumes the TCP request and makes a best-effort attempt to enqueue the RFC 9293 reset without waiting for outbound capacity. Local output congestion may discard the reset without error. It reports syscall.EADDRNOTAVAIL when the intercepted destination is no longer admitted and syscall.ENETUNREACH when no return route remains; the rejection decision remains terminal.

type TCPHeaderOption

type TCPHeaderOption struct {
	// Kind is the TCP option kind.
	Kind uint8
	// Data is the option value following Kind and Length.
	Data []byte
}

TCPHeaderOption is one TCP option in semantic wire order. Data excludes the Kind and Length bytes. End and NOP require empty Data; every other kind is encoded with a Length byte. HeaderOptions returns Data slices that borrow TCPSegment.Options, while SetHeaderOptions copies every Data slice.

func (TCPHeaderOption) IsSACKPermitted

func (o TCPHeaderOption) IsSACKPermitted() bool

IsSACKPermitted reports whether o is a well-formed SACK-Permitted option.

func (TCPHeaderOption) MaximumSegmentSize

func (o TCPHeaderOption) MaximumSegmentSize() (uint16, bool)

MaximumSegmentSize returns the raw MSS value when o is a well-formed MSS option. It does not replace zero or clamp the value to a path limit.

func (TCPHeaderOption) SACKBlocks

func (o TCPHeaderOption) SACKBlocks() ([]TCPSACKBlock, bool)

SACKBlocks returns the raw blocks when o is a well-formed SACK option. It does not classify DSACK, filter blocks against a send window, or merge them.

func (*TCPHeaderOption) SetMaximumSegmentSize

func (o *TCPHeaderOption) SetMaximumSegmentSize(value uint16)

SetMaximumSegmentSize replaces o with an MSS option containing value.

func (*TCPHeaderOption) SetSACKBlocks

func (o *TCPHeaderOption) SetSACKBlocks(blocks []TCPSACKBlock) error

SetSACKBlocks replaces o with a SACK option containing one to four blocks. It copies the block values and leaves o unchanged when the count is invalid.

func (*TCPHeaderOption) SetSACKPermitted

func (o *TCPHeaderOption) SetSACKPermitted()

SetSACKPermitted replaces o with a SACK-Permitted option.

func (*TCPHeaderOption) SetTimestamp

func (o *TCPHeaderOption) SetTimestamp(value, echo uint32)

SetTimestamp replaces o with a Timestamp option containing value and echo.

func (*TCPHeaderOption) SetWindowScale

func (o *TCPHeaderOption) SetWindowScale(value uint8)

SetWindowScale replaces o with a Window Scale option containing value.

func (TCPHeaderOption) Timestamp

func (o TCPHeaderOption) Timestamp() (value, echo uint32, ok bool)

Timestamp returns the raw TSval and TSecr values when o is a well-formed Timestamp option.

func (TCPHeaderOption) WindowScale

func (o TCPHeaderOption) WindowScale() (uint8, bool)

WindowScale returns the raw scale when o is a well-formed Window Scale option. Values above RFC 7323's operational maximum are not clamped.

type TCPListener

type TCPListener struct {
	// contains filtered or unexported fields
}

TCPListener is a passive userspace TCP endpoint.

func (*TCPListener) Accept

func (l *TCPListener) Accept() (net.Conn, error)

Accept waits for and returns the next completed passive connection.

func (*TCPListener) Addr

func (l *TCPListener) Addr() net.Addr

Addr returns the bound TCP endpoint.

func (*TCPListener) Close

func (l *TCPListener) Close() error

Close stops listening without closing connections already returned by Accept.

func (*TCPListener) Info

func (l *TCPListener) Info() TCPListenerInfo

Info returns queue pressure, handshake outcomes, and SYN-cookie activity for this listener. The final snapshot remains available after Close.

func (*TCPListener) SetDeadline

func (l *TCPListener) SetDeadline(deadline time.Time) error

SetDeadline sets the deadline for subsequent Accept calls.

type TCPListenerInfo

type TCPListenerInfo struct {
	// LocalAddress is the bound listener endpoint.
	LocalAddress netip.AddrPort
	// Closed reports whether the listener was closed when sampled.
	Closed bool
	// AcceptQueueConnections is the number of completed connections awaiting Accept.
	AcceptQueueConnections int
	// AcceptQueueCapacity is the completed-connection queue limit.
	AcceptQueueCapacity int
	// AcceptQueuePeak is the lifetime peak completed-connection queue depth.
	AcceptQueuePeak int
	// SYNBacklogConnections is the number of stateful handshakes in progress.
	SYNBacklogConnections int
	// SYNBacklogCapacity is the stateful handshake limit.
	SYNBacklogCapacity int
	// SYNBacklogPeak is the lifetime peak number of stateful handshakes.
	SYNBacklogPeak int
	// SYNsReceived counts valid initial SYN segments dispatched to the listener.
	SYNsReceived uint64
	// StatefulHandshakes counts connection states allocated for initial SYNs.
	StatefulHandshakes uint64
	// HandshakeCompletions counts passive handshakes that reached the accept queue.
	HandshakeCompletions uint64
	// HandshakeFailures counts stateful passive handshakes that did not complete.
	HandshakeFailures uint64
	// HandshakeTimeouts counts stateful passive handshakes that timed out.
	HandshakeTimeouts uint64
	// SYNCookiesSent counts stateless SYN-cookie responses.
	SYNCookiesSent uint64
	// SYNCookiesAccepted counts valid cookie acknowledgements.
	SYNCookiesAccepted uint64
	// SYNCookiesRejected counts invalid or stale cookie acknowledgements.
	SYNCookiesRejected uint64
	// AcceptQueueDrops counts completed handshakes rejected by a full accept queue.
	AcceptQueueDrops uint64
	// AcceptedConnections counts connections returned successfully by Accept.
	AcceptedConnections uint64
}

TCPListenerInfo is a point-in-time diagnostic snapshot of one passive TCP endpoint. Queue peaks and counters cover the listener's complete lifetime.

type TCPSACKBlock

type TCPSACKBlock struct {
	// LeftEdge is the sequence number of the first acknowledged byte.
	LeftEdge uint32
	// RightEdge is the sequence number immediately after the acknowledged range.
	RightEdge uint32
}

TCPSACKBlock is one half-open sequence range carried by a TCP SACK option. Edges use TCP's wrapping 32-bit sequence space; interpreting their order requires connection context that the standalone wire codec does not have.

type TCPSegment

type TCPSegment struct {
	// Source is the source IP address and TCP port.
	Source netip.AddrPort
	// Destination is the destination IP address and TCP port.
	Destination netip.AddrPort
	// SequenceNumber is the first sequence number represented by the segment.
	SequenceNumber uint32
	// AcknowledgmentNumber is meaningful when TCPFlagACK is set.
	AcknowledgmentNumber uint32
	// Flags contains the eight current TCP control flags and the historic NS bit.
	Flags uint16
	// WindowSize is the unscaled advertised receive window.
	WindowSize uint16
	// UrgentPointer is meaningful when TCPFlagURG is set.
	UrgentPointer uint16
	// Options contains the exact parsed option area, including padding.
	// Construction also accepts an unpadded option sequence and adds zero padding;
	// bytes following an End of Option List are normalized to zero as RFC 9293
	// requires for generated segments.
	Options []byte
	// Payload is the segment's application data.
	Payload []byte
}

TCPSegment is the semantic representation of one checksummed TCP segment. Source and Destination provide both wire ports and the IP pseudo-header addresses. IPPacket.TCPSegment borrows Options and Payload from the packet; callers must replace or copy those slices before modifying unowned input. Construction normalizes IPv4-mapped IPv6 addresses to IPv4.

func (TCPSegment) AppendBinary

func (s TCPSegment) AppendBinary(dst []byte) ([]byte, error)

AppendBinary appends the complete TCP segment wire encoding to dst. Source and Destination contribute to its pseudo-header checksum but are not themselves encoded. It validates every field before changing dst, does not retain any input slice, and permits the destination to share backing storage with Options or Payload. On validation failure it returns the original dst unchanged.

func (TCPSegment) HeaderOptions

func (s TCPSegment) HeaderOptions() ([]TCPHeaderOption, error)

HeaderOptions parses the exact TCP option sequence. The returned slice owns its option descriptors, but each Data field borrows Options. End is returned as the final descriptor and bytes after it are ignored as receiver padding. Recognized kinds with nonstandard lengths remain available as raw options; their typed accessors report ok=false.

func (TCPSegment) MarshalBinary

func (s TCPSegment) MarshalBinary() ([]byte, error)

MarshalBinary returns the complete TCP segment wire encoding. Source and Destination contribute to its pseudo-header checksum but are not themselves encoded. MarshalBinary is semantically identical to AppendBinary(nil).

func (*TCPSegment) SetHeaderOptions

func (s *TCPSegment) SetHeaderOptions(options []TCPHeaderOption) error

SetHeaderOptions replaces Options with the encoded option sequence. It preserves unknown kinds, duplicates, and order, copies all input data, and leaves s unchanged on failure. An End option must be last. The final TCP header padding is added by MarshalBinary or AppendBinary.

type TCPSocketDefaults

type TCPSocketDefaults struct {
	// CongestionControl selects a registered algorithm by name for new
	// connections. The zero value selects CUBIC. UpdateConfig also applies a
	// changed value to established connections without an explicit
	// per-connection override. It must be empty when CongestionControlFactory is
	// set.
	CongestionControl string
	// CongestionControlFactory selects an immutable local factory without
	// process-wide registration. It must be created by
	// NewCongestionControlFactory and is mutually exclusive with
	// CongestionControl. The factory creates an independent controller for every
	// connection and may be shared safely by multiple stacks and listeners.
	CongestionControlFactory *CongestionControlFactory
	// ReceiveBuffer is the initial application receive capacity.
	ReceiveBuffer int
	// MaximumReceiveBuffer bounds automatic receive tuning.
	MaximumReceiveBuffer int
	// SendBuffer is the initial application send capacity.
	SendBuffer int
	// MaximumSendBuffer bounds automatic send tuning.
	MaximumSendBuffer int
	// MaximumPacingRate caps the paced-data rate of new TCP connections in
	// bytes per second. Zero leaves the pacing rate unlimited. Initial and
	// control bursts mean this is not a strict byte-rate shaper.
	MaximumPacingRate uint64
	// AcceptQueue bounds completed connections waiting for Accept.
	AcceptQueue int
	// SYNBacklog bounds stateful handshakes before SYN cookies are used.
	SYNBacklog int
	// KeepAlive enables keepalive probes on new connections.
	KeepAlive bool
	// KeepAliveConfig supplies the default probe timing and retry count.
	KeepAliveConfig KeepAliveConfig
	// IdleTimeout closes a connection after receive inactivity. Zero disables it.
	IdleTimeout time.Duration
	// UserTimeout bounds how long transmitted data may remain unacknowledged,
	// or buffered data may remain unsent behind a zero window. Zero disables
	// this custom bound while retaining the normal TCP retry limits.
	UserTimeout time.Duration
	// DisableNoDelay makes new connections start with Nagle coalescing enabled.
	DisableNoDelay bool
	// TrafficClass supplies IPv4 TOS or IPv6 Traffic Class DSCP bits. TCP
	// controls the two ECN bits independently.
	TrafficClass uint8
	// FlowLabel fixes the IPv6 Flow Label on new connections. Zero selects a
	// stable RFC 6437-style label derived from the connection tuple.
	FlowLabel uint32
}

TCPSocketDefaults configures policies inherited by newly created TCP connections and listeners. Zero fields retain the package defaults.

type TCPState

type TCPState uint8

TCPState identifies the current RFC 9293 connection state exposed by TCPConn.Info.

const (
	// TCPStateClosed indicates that no connection state remains.
	TCPStateClosed TCPState = iota
	// TCPStateSYNReceived indicates a passive handshake awaiting its final ACK.
	TCPStateSYNReceived
	// TCPStateSYNSent indicates an active handshake awaiting a SYN-ACK.
	TCPStateSYNSent
	// TCPStateEstablished indicates bidirectional data transfer state.
	TCPStateEstablished
	// TCPStateFINWait1 indicates that the local FIN is not yet acknowledged.
	TCPStateFINWait1
	// TCPStateFINWait2 indicates that the local FIN is acknowledged while the peer remains open.
	TCPStateFINWait2
	// TCPStateCloseWait indicates that the peer closed first.
	TCPStateCloseWait
	// TCPStateClosing indicates simultaneous close awaiting the local FIN ACK.
	TCPStateClosing
	// TCPStateLastACK indicates a locally sent FIN after the peer closed first.
	TCPStateLastACK
	// TCPStateTimeWait indicates retained state after an active or simultaneous close.
	TCPStateTimeWait
)

func (TCPState) String

func (s TCPState) String() string

String returns the conventional RFC 9293 state name.

type UDPConn

type UDPConn struct {
	// contains filtered or unexported fields
}

UDPConn is a connected or unconnected userspace UDP socket.

func (*UDPConn) Broadcast

func (c *UDPConn) Broadcast() (bool, error)

Broadcast reports the SO_BROADCAST-equivalent output permission.

func (*UDPConn) Close

func (c *UDPConn) Close() error

Close unregisters the socket and wakes blocked reads.

func (*UDPConn) ConfirmPathMTU

func (c *UDPConn) ConfirmPathMTU(mtu int) error

ConfirmPathMTU records application-level acknowledgement of a connected UDP probe. mtu is the complete IP packet size, not the UDP payload size.

func (*UDPConn) ConfirmPathMTUFor

func (c *UDPConn) ConfirmPathMTUFor(target netip.Addr, mtu int) error

ConfirmPathMTUFor is the unconnected form of ConfirmPathMTU.

func (*UDPConn) ExcludeSourceSpecificGroup

func (c *UDPConn) ExcludeSourceSpecificGroup(group, source netip.Addr) error

ExcludeSourceSpecificGroup blocks source on an existing any-source membership. It returns EINVAL for an SSM group.

func (*UDPConn) IncludeSourceSpecificGroup

func (c *UDPConn) IncludeSourceSpecificGroup(group, source netip.Addr) error

IncludeSourceSpecificGroup removes a source block from an existing any-source membership. It returns EINVAL for an SSM group.

func (*UDPConn) Info

func (c *UDPConn) Info() UDPConnInfo

Info returns a diagnostic snapshot of the socket and its receive queue.

func (*UDPConn) JoinGroup

func (c *UDPConn) JoinGroup(group netip.Addr) error

JoinGroup joins an any-source multicast group. It is the single-interface equivalent of MCAST_JOIN_GROUP and x/net's JoinGroup. RFC 4604 SSM groups return EINVAL and require JoinSourceSpecificGroup.

func (*UDPConn) JoinSourceSpecificGroup

func (c *UDPConn) JoinSourceSpecificGroup(group, source netip.Addr) error

JoinSourceSpecificGroup adds source to an INCLUDE-mode membership, creating the membership when this is its first source.

func (*UDPConn) LeaveGroup

func (c *UDPConn) LeaveGroup(group netip.Addr) error

LeaveGroup leaves a multicast group regardless of its source-filter mode.

func (*UDPConn) LeaveSourceSpecificGroup

func (c *UDPConn) LeaveSourceSpecificGroup(group, source netip.Addr) error

LeaveSourceSpecificGroup removes source from an INCLUDE-mode membership. It leaves the group when source was its final entry.

func (*UDPConn) LocalAddr

func (c *UDPConn) LocalAddr() net.Addr

LocalAddr returns the unspecified family address and allocated port.

func (*UDPConn) MulticastHopLimit

func (c *UDPConn) MulticastHopLimit() (int, error)

MulticastHopLimit returns the IPv4 multicast TTL or IPv6 multicast Hop Limit used by subsequent writes.

func (*UDPConn) MulticastLoopback

func (c *UDPConn) MulticastLoopback() (bool, error)

MulticastLoopback reports whether transmitted multicast packets are copied to matching local memberships.

func (*UDPConn) MulticastSourceFilter

func (c *UDPConn) MulticastSourceFilter(group netip.Addr) (MulticastSourceFilter, error)

MulticastSourceFilter returns the complete source policy for group.

func (*UDPConn) PathMTUDiscovery

func (c *UDPConn) PathMTUDiscovery() (PathMTUDiscovery, error)

PathMTUDiscovery returns the Linux-compatible IP_MTU_DISCOVER policy used by subsequent UDP writes.

func (*UDPConn) Read

func (c *UDPConn) Read(buffer []byte) (int, error)

Read receives the next datagram from a connected remote endpoint.

func (*UDPConn) ReadBatch

func (c *UDPConn) ReadBatch(messages []SocketMessage, flags int) (int, error)

ReadBatch reads one or more UDP messages using the SocketMessage layout shared by x/net/ipv4 and x/net/ipv6. The first message follows the socket's blocking and deadline semantics; after it succeeds, the method drains only messages already queued. MessageFlagDontWait also makes the first read nonblocking.

func (*UDPConn) ReadError

func (c *UDPConn) ReadError() (*net.OpError, error)

ReadError returns the oldest queued asynchronous network error without blocking. An empty queue reports EAGAIN, like a Linux MSG_ERRQUEUE read on a nonblocking descriptor. Ordinary operations may consume the pending socket error without removing this entry from the extended error queue.

func (*UDPConn) ReadFrom

func (c *UDPConn) ReadFrom(buffer []byte) (int, net.Addr, error)

ReadFrom returns the next complete datagram or socket error.

func (*UDPConn) ReadFromUDP

func (c *UDPConn) ReadFromUDP(buffer []byte) (int, *net.UDPAddr, error)

ReadFromUDP acts like ReadFrom but returns a UDPAddr.

func (*UDPConn) ReadFromUDPAddrPort

func (c *UDPConn) ReadFromUDPAddrPort(buffer []byte) (int, netip.AddrPort, error)

ReadFromUDPAddrPort acts like ReadFrom but returns a netip.AddrPort.

func (*UDPConn) ReadFromUDPAddrPortWithBuffer

func (c *UDPConn) ReadFromUDPAddrPortWithBuffer(getBuffer func(sizeHint int) []byte) (int, netip.AddrPort, error)

ReadFromUDPAddrPortWithBuffer is the netip.AddrPort form of ReadFromWithBuffer.

It has the same callback, buffer ownership, truncation, and nil-callback semantics as ReadFromWithBuffer.

This is an experimental API and is not covered by the package's stability guarantees.

func (*UDPConn) ReadFromUDPWithBuffer

func (c *UDPConn) ReadFromUDPWithBuffer(getBuffer func(sizeHint int) []byte) (int, *net.UDPAddr, error)

ReadFromUDPWithBuffer is the *net.UDPAddr form of ReadFromWithBuffer.

It has the same callback, buffer ownership, truncation, and nil-callback semantics as ReadFromWithBuffer.

This is an experimental API and is not covered by the package's stability guarantees.

func (*UDPConn) ReadFromWithBuffer

func (c *UDPConn) ReadFromWithBuffer(getBuffer func(sizeHint int) []byte) (int, net.Addr, error)

ReadFromWithBuffer reads the next complete datagram like ReadFrom, obtaining the destination buffer lazily from getBuffer and returning its source address.

If a datagram is available, getBuffer is called once after it is dequeued. It is not called when the operation returns before a datagram is available because of an error or deadline. The callback receives the complete UDP payload length as an advisory size hint. It runs without c's connection state lock and must return promptly; it must not call a read method on c. The caller owns the returned slice, and c does not retain it. A nil callback returns EINVAL. A short returned slice truncates and consumes the datagram, matching ReadFrom.

This is an experimental API and is not covered by the package's stability guarantees.

func (*UDPConn) ReadMsgUDP

func (c *UDPConn) ReadMsgUDP(buffer, oob []byte) (n, oobn, flags int, address *net.UDPAddr, err error)

ReadMsgUDP reads one datagram and Linux-compatible packet-info ancillary data. The control message identifies the packet's IP destination address.

func (*UDPConn) ReadMsgUDPAddrPort

func (c *UDPConn) ReadMsgUDPAddrPort(buffer, oob []byte) (n, oobn, flags int, source netip.AddrPort, err error)

ReadMsgUDPAddrPort is the netip.AddrPort form of ReadMsgUDP.

func (*UDPConn) ReadWithBuffer

func (c *UDPConn) ReadWithBuffer(getBuffer func(sizeHint int) []byte) (int, error)

ReadWithBuffer reads the next datagram from a connected remote endpoint like Read, obtaining the destination buffer lazily from getBuffer.

If a datagram is available, getBuffer is called once after it is dequeued. It is not called when the operation returns before a datagram is available because of an error or deadline. The callback receives the complete UDP payload length as an advisory size hint. It runs without c's connection state lock and must return promptly; it must not call a read method on c. The caller owns the returned slice, and c does not retain it. A nil callback returns EINVAL. A short returned slice truncates and consumes the datagram, matching Read.

This is an experimental API and is not covered by the package's stability guarantees.

func (*UDPConn) ReceiveErrors

func (c *UDPConn) ReceiveErrors() (bool, error)

ReceiveErrors reports whether asynchronous errors are retained for ReadError. When disabled, unconnected sockets do not report those errors; connected sockets report hard errors on ordinary reads or writes. It also reports whether immediate failure to admit unicast or external-link non-unicast output is reported as ENOBUFS.

func (*UDPConn) RemoteAddr

func (c *UDPConn) RemoteAddr() net.Addr

RemoteAddr returns the connected UDP endpoint, or nil for an unconnected packet socket.

func (*UDPConn) SetBroadcast

func (c *UDPConn) SetBroadcast(enabled bool) error

SetBroadcast changes the SO_BROADCAST-equivalent output permission.

func (*UDPConn) SetDeadline

func (c *UDPConn) SetDeadline(deadline time.Time) error

SetDeadline updates both read and write deadlines.

func (*UDPConn) SetFlowLabel

func (c *UDPConn) SetFlowLabel(label uint32) error

SetFlowLabel changes the default IPv6 Flow Label. Zero explicitly disables automatic labeling for this socket.

func (*UDPConn) SetHopLimit

func (c *UDPConn) SetHopLimit(hopLimit int) error

SetHopLimit changes the default IPv4 TTL or IPv6 Hop Limit for subsequent writes. Zero is valid only on a dedicated IPv6 socket; it is ambiguous on a dual-stack socket because IPv4 TTL zero is invalid. Per-packet message control data may override the value.

func (*UDPConn) SetMulticastHopLimit

func (c *UDPConn) SetMulticastHopLimit(hopLimit int) error

SetMulticastHopLimit changes the IPv4 multicast TTL or IPv6 multicast Hop Limit. Zero confines output to this host.

func (*UDPConn) SetMulticastLoopback

func (c *UDPConn) SetMulticastLoopback(enabled bool) error

SetMulticastLoopback controls delivery of transmitted multicast packets to matching local memberships.

func (*UDPConn) SetMulticastSourceFilter

func (c *UDPConn) SetMulticastSourceFilter(group netip.Addr, filter MulticastSourceFilter) error

SetMulticastSourceFilter atomically replaces the complete INCLUDE/EXCLUDE source policy for a previously joined group. An empty INCLUDE filter leaves that membership, matching Linux MCAST_MSFILTER behavior. EXCLUDE returns EINVAL for an RFC 4604 SSM group.

func (*UDPConn) SetPathMTUDiscovery

func (c *UDPConn) SetPathMTUDiscovery(mode PathMTUDiscovery) error

SetPathMTUDiscovery changes the Linux-compatible IP_MTU_DISCOVER policy for subsequent UDP writes.

func (*UDPConn) SetReadBuffer

func (c *UDPConn) SetReadBuffer(bytes int) error

SetReadBuffer changes the approximate memory capacity shared by the datagram and asynchronous-error receive queues. Existing entries are retained when the capacity shrinks; later arrivals are dropped until enough space becomes available.

func (*UDPConn) SetReadDeadline

func (c *UDPConn) SetReadDeadline(deadline time.Time) error

SetReadDeadline updates the next ReadFrom deadline.

func (*UDPConn) SetReceiveErrors

func (c *UDPConn) SetReceiveErrors(enabled bool) error

SetReceiveErrors controls whether asynchronous network errors are retained for ReadError. It also makes a write fail with ENOBUFS when immediate admission of unicast output or the external-link copy of multicast or broadcast output fails. It does not report packets displaced after admission. Receive-side non-unicast loopback copies remain best effort. By default, unconnected sockets do not report asynchronous ICMP errors, although correlated PMTU updates still apply. Connected sockets report hard errors from ordinary reads or writes before queued datagrams. When enabled, eligible soft errors also reach ordinary operations and ReadError. Disabling the option clears the extended queue but preserves a pending ordinary error. When disabled, immediate output admission failures are silent.

func (*UDPConn) SetTrafficClass

func (c *UDPConn) SetTrafficClass(value int) error

SetTrafficClass changes the default IPv4 TOS or IPv6 Traffic Class byte.

func (*UDPConn) SetWriteBuffer

func (c *UDPConn) SetWriteBuffer(bytes int) error

SetWriteBuffer validates the standard socket option but otherwise has no work to do: UDP writes make one immediate bounded link-queue admission attempt and retain no per-socket transmit buffer to resize.

func (*UDPConn) SetWriteDeadline

func (c *UDPConn) SetWriteDeadline(deadline time.Time) error

SetWriteDeadline sets the deadline checked before future writes.

func (*UDPConn) Write

func (c *UDPConn) Write(payload []byte) (int, error)

Write sends one datagram to the connected remote endpoint.

func (*UDPConn) WriteBatch

func (c *UDPConn) WriteBatch(messages []SocketMessage, flags int) (int, error)

WriteBatch writes a prefix of UDP messages using scatter/gather payloads. MessageFlagDontWait is accepted for Linux compatibility; device admission is already nonblocking for every datagram write. Other flags are unsupported.

func (*UDPConn) WriteMsgUDP

func (c *UDPConn) WriteMsgUDP(payload, oob []byte, address *net.UDPAddr) (n, oobn int, err error)

WriteMsgUDP writes a payload using Linux-compatible packet-info ancillary data. A connected socket requires a nil address.

func (*UDPConn) WriteMsgUDPAddrPort

func (c *UDPConn) WriteMsgUDPAddrPort(payload, oob []byte, address netip.AddrPort) (n, oobn int, err error)

WriteMsgUDPAddrPort is the netip.AddrPort form of WriteMsgUDP. A connected socket requires an invalid address.

func (*UDPConn) WritePathMTUProbe

func (c *UDPConn) WritePathMTUProbe(payload []byte) (int, error)

WritePathMTUProbe sends one connected UDP datagram without IPv4 or IPv6 fragmentation, permitting a size above the confirmed PMTU up to the configured first-hop MTU. The application must confirm delivery separately.

func (*UDPConn) WritePathMTUProbeTo

func (c *UDPConn) WritePathMTUProbeTo(payload []byte, target netip.AddrPort) (int, error)

WritePathMTUProbeTo is the unconnected netip form of WritePathMTUProbe.

func (*UDPConn) WriteTo

func (c *UDPConn) WriteTo(payload []byte, address net.Addr) (int, error)

WriteTo sends one datagram, fragmenting its IP payload when required.

func (*UDPConn) WriteToUDP

func (c *UDPConn) WriteToUDP(payload []byte, address *net.UDPAddr) (int, error)

WriteToUDP acts like WriteTo but accepts a UDPAddr directly.

func (*UDPConn) WriteToUDPAddrPort

func (c *UDPConn) WriteToUDPAddrPort(payload []byte, address netip.AddrPort) (int, error)

WriteToUDPAddrPort acts like WriteTo but accepts a netip.AddrPort directly.

type UDPConnInfo

type UDPConnInfo struct {
	// LocalAddress is the bound local endpoint; an unspecified address denotes
	// a wildcard binding.
	LocalAddress netip.AddrPort
	// RemoteAddress is the connected peer, or an invalid endpoint for an
	// unconnected socket.
	RemoteAddress netip.AddrPort
	// Closed reports whether the socket was closed when the snapshot was taken.
	Closed bool
	// ReceiveQueuePackets is the number of complete datagrams awaiting a read.
	ReceiveQueuePackets int
	// ReceiveQueueBytes is the accounted payload and metadata retained by the
	// receive queue.
	ReceiveQueueBytes int
	// ReceiveQueueCapacity is the configured accounting-byte limit of the
	// combined datagram and extended-error queues, not an exact heap limit.
	ReceiveQueueCapacity int
	// ReceiveErrors reports whether asynchronous network errors are retained
	// for ReadError. Otherwise, unconnected sockets do not report those errors;
	// connected sockets report hard errors on ordinary reads or writes. It also
	// reports whether immediate failure to admit unicast or external-link
	// non-unicast output is reported as ENOBUFS.
	ReceiveErrors bool
	// ErrorQueueEntries is the number of extended errors awaiting ReadError.
	ErrorQueueEntries int
	// ErrorQueueBytes is the accounted metadata and quoted packet data retained
	// by the asynchronous error queue.
	ErrorQueueBytes int
	// ErrorsDropped counts extended-error entries discarded because the
	// configured receive-buffer budget was exhausted.
	ErrorsDropped uint64
	// PacketsSent counts successful UDP socket write results. It includes writes
	// silently lost during bounded output admission under the default
	// ReceiveErrors policy and remains cumulative if bounded link scheduling
	// later drops a packet.
	PacketsSent uint64
	// BytesSent counts payload bytes represented by those successful writes.
	BytesSent uint64
	// PacketsReceived counts datagrams accepted into the receive queue.
	PacketsReceived uint64
	// BytesReceived counts UDP payload bytes accepted into the receive queue.
	BytesReceived uint64
	// PacketsDropped counts datagrams rejected because the socket was closed or
	// its receive queue lacked capacity.
	PacketsDropped uint64
	// ICMPErrors counts matching asynchronous ICMP errors delivered to the
	// socket.
	ICMPErrors uint64
	// PathMTU is the complete-IP-packet PMTU for a connected unicast peer, or
	// zero when no such path exists.
	PathMTU int
	// PathMTUDiscovery is the Linux-compatible source-fragmentation and PMTU
	// policy used by subsequent writes.
	PathMTUDiscovery PathMTUDiscovery
	// HopLimit is the default unicast IPv4 TTL or IPv6 Hop Limit.
	HopLimit int
	// MulticastHopLimit is the default multicast IPv4 TTL or IPv6 Hop Limit.
	MulticastHopLimit int
	// MulticastLoopback reports whether transmitted multicast is delivered to
	// matching local memberships.
	MulticastLoopback bool
	// Broadcast reports whether IPv4 broadcast output is permitted.
	Broadcast bool
	// TrafficClass is the default IPv4 TOS or IPv6 Traffic Class byte.
	TrafficClass uint8
	// FlowLabel is the effective IPv6 Flow Label; it is zero for IPv4 sockets.
	FlowLabel uint32
	// LastError is the most recently recorded socket operation or asynchronous
	// network error.
	LastError error
}

UDPConnInfo is a point-in-time diagnostic snapshot of one UDP socket. Traffic byte counters measure UDP payload; receive-queue byte values also include the stack's per-datagram accounting overhead.

type UDPDatagram

type UDPDatagram struct {
	// Source is the source IP address and UDP port.
	Source netip.AddrPort
	// Destination is the destination IP address and UDP port.
	Destination netip.AddrPort
	// ChecksumDisabled requests the optional zero IPv4 UDP checksum. It is
	// invalid for IPv6.
	ChecksumDisabled bool
	// Payload is the datagram body.
	Payload []byte
}

UDPDatagram is the semantic representation of one UDP datagram. Source and Destination provide both wire ports and the IP pseudo-header addresses. IPPacket.UDPDatagram borrows Payload from the packet; callers must replace or copy it before modifying unowned input. Construction normalizes IPv4-mapped IPv6 addresses to IPv4.

func (UDPDatagram) AppendBinary

func (d UDPDatagram) AppendBinary(dst []byte) ([]byte, error)

AppendBinary appends the complete UDP datagram wire encoding to dst. Source and Destination contribute to its pseudo-header checksum but only their ports are encoded. It validates every field before changing dst, does not retain any input slice, and permits the destination to share backing storage with Payload. On validation failure it returns the original dst unchanged.

func (UDPDatagram) MarshalBinary

func (d UDPDatagram) MarshalBinary() ([]byte, error)

MarshalBinary returns the complete UDP datagram wire encoding. Source and Destination contribute to its pseudo-header checksum but only their ports are encoded. MarshalBinary is semantically identical to AppendBinary(nil).

type UDPForwarder

type UDPForwarder struct {
	// contains filtered or unexported fields
}

UDPForwarder owns the single fallback UDP handler installed on a stack. It may be installed before or after Stack.Start.

func NewUDPForwarder

func NewUDPForwarder(stack *Stack, options UDPForwarderOptions, handler UDPForwarderHandler) (*UDPForwarder, error)

NewUDPForwarder installs a fallback handler for otherwise unhandled UDP datagrams. Only one UDP forwarder may be active per stack. Promiscuous mode is not required for unhandled traffic addressed to LocalAddresses; Config.Promiscuous is required only for nonlocal destination addresses. Installing a forwarder does not start the stack.

func (*UDPForwarder) Close

func (f *UDPForwarder) Close() error

Close removes the UDP fallback handler and invalidates undecided requests. It does not wait for running handlers to return. An action already claimed by a handler or detached responder may finish, and accepted UDP endpoints remain open.

func (UDPForwarder) Done

func (f UDPForwarder) Done() <-chan struct{}

Done is closed when the forwarder is closed directly or by Stack.Close.

func (*UDPForwarder) Info

func (f *UDPForwarder) Info() ForwarderInfo

Info returns one UDP forwarder diagnostic snapshot.

type UDPForwarderHandler

type UDPForwarderHandler func(*UDPForwarderRequest)

UDPForwarderHandler decides the fate of one valid, otherwise unhandled UDP datagram. MIPS calls it synchronously from Stack.Write or the loopback worker. It must return promptly and must not wait for traffic whose delivery depends on the blocked call. Concurrent Stack.Write calls may invoke the handler concurrently, so shared state must be synchronized. While one request is undecided, concurrent datagrams for the same four-tuple are dropped. The handler must call Accept, Listen, Detach, DetachForReplies, Drop, Reject, or at least one Reply before returning; returning without an action drops the datagram. Reply may be repeated and does not prevent a later terminal action. Except for Detach and DetachForReplies, every action and Reply call must finish before the handler returns. The request and its Payload must not be retained after that point, but a UDPConn returned by Accept or Listen and a responder returned by either detach method may outlive the callback. Responder output remains subject to the originating forwarder's state. The initial datagram remains subject to the returned UDPConn's configured receive capacity.

type UDPForwarderOptions

type UDPForwarderOptions struct{}

UDPForwarderOptions reserves UDP interception policy for future extension.

type UDPForwarderRequest

type UDPForwarderRequest struct {
	// contains filtered or unexported fields
}

UDPForwarderRequest is one valid datagram that did not match an ordinary or previously forwarded UDP endpoint. The handler may reply repeatedly before selecting at most one terminal action. Payload is valid only until the handler returns; Accept and Listen offer a copy to the returned UDPConn's capacity-bounded receive queue.

func (*UDPForwarderRequest) Accept

func (r *UDPForwarderRequest) Accept(options ...SocketOption) (*UDPConn, error)

Accept creates a connected UDP endpoint, offers a copy of the triggering datagram to its receive queue, and registers the complete intercepted four-tuple for future delivery. It does not wait for remote traffic. The returned UDPConn is bound to Destination and connected to Source: Read receives only that source, Write replies to it, and destination-taking methods such as WriteTo return net.ErrWriteToConnected. The endpoint remains open if the forwarder closes. The handler must wait for Accept to return before returning itself, but may then retain the connection or hand it to another goroutine. The triggering datagram may be dropped when it exceeds the configured receive capacity; later datagrams still use the registered endpoint. Accept consumes the request even when endpoint creation returns an error and may be called after any number of Reply or ReplyFrom attempts. Options are validated before the request is claimed; an invalid option leaves it available for another action, and the option slice is not retained.

func (*UDPForwarderRequest) Detach

Detach transfers one UDP request out of the synchronous handler lifetime. On success it removes the request from the forwarder's pending set and returns a caller-owned flow and payload snapshot. The responder points back to the originating forwarder's state for output and Done, but the forwarder does not retain the responder or impose a capacity or timeout. The caller may hand it to another goroutine or discard it without a terminal action. Detach itself is the request's action and consumes the request even when it returns an error. It may be called after any number of Reply or ReplyFrom attempts; the responder remains available for further replies.

func (*UDPForwarderRequest) DetachForReplies

func (r *UDPForwarderRequest) DetachForReplies() (*UDPForwarderResponder, error)

DetachForReplies transfers the UDP request into an asynchronous responder that retains only Flow, reply operations, and Done. Payload returns nil, and Reject and Drop report net.ErrClosed. Unlike Detach, it does not copy the triggering payload or retain an ICMP rejection quote. The caller owns the responder and may discard it without a terminal action. RestrictToReplies is an idempotent no-op on the returned responder.

func (*UDPForwarderRequest) Drop

func (r *UDPForwarderRequest) Drop() error

Drop consumes the UDP datagram without packet I/O. It may follow any number of Reply or ReplyFrom attempts, may wait briefly for forwarder bookkeeping, and does not wait for the network or output queue.

func (*UDPForwarderRequest) Flow

Flow returns the original inbound UDP four-tuple. The method may be called only during the handler, but the returned value is an independent copy that may be retained.

func (*UDPForwarderRequest) Listen

func (r *UDPForwarderRequest) Listen(options ...SocketOption) (*UDPConn, error)

Listen creates an unconnected UDP endpoint bound to Destination, offers a copy of the triggering datagram to its receive queue, and registers that local endpoint for future datagrams from any source. It does not wait for remote traffic. ReadFrom reports each source and WriteTo may address different peers. The destination must not already have an ordinary binding or an accepted forwarded flow; such an ownership conflict reports syscall.EADDRINUSE. The endpoint remains open if the forwarder closes. The handler must wait for Listen to return before returning itself, but may then retain the connection or hand it to another goroutine. The triggering datagram may be dropped when it exceeds the configured receive capacity; later datagrams still use the registered endpoint. Listen consumes the request even when endpoint creation returns an error and may be called after any number of Reply or ReplyFrom attempts. Options are validated before the request is claimed; an invalid option leaves it available for another action, and the option slice is not retained.

func (*UDPForwarderRequest) Payload

func (r *UDPForwarderRequest) Payload() []byte

Payload returns the triggering UDP payload. The returned slice aliases packet-delivery storage, must not be modified, and is valid only until the handler returns.

func (*UDPForwarderRequest) Reject

func (r *UDPForwarderRequest) Reject() error

Reject consumes the UDP datagram and makes a best-effort attempt to enqueue ICMP Port Unreachable without waiting for outbound capacity. Local output congestion may discard the response without error. It reports syscall.EADDRNOTAVAIL when the intercepted destination is no longer admitted and syscall.ENETUNREACH when no return route remains. It may follow any number of Reply or ReplyFrom attempts; the rejection decision remains terminal.

func (*UDPForwarderRequest) Reply

func (r *UDPForwarderRequest) Reply(payload []byte) (int, error)

Reply sends one reverse-flow datagram from Destination to Source without retaining a UDP endpoint. Use ReplyFrom to select a different source. The method may be called repeatedly or concurrently, including before a later terminal action, but every call must finish before the handler returns. Each call uses the current Config.UDP output defaults and makes one immediate best-effort output attempt. Local output congestion may discard the datagram or any of its source fragments without error. Other errors may be retried.

func (*UDPForwarderRequest) ReplyFrom

func (r *UDPForwarderRequest) ReplyFrom(payload []byte, source netip.AddrPort) (int, error)

ReplyFrom sends one datagram to Flow().Source using source as its IP address and UDP port, without retaining an endpoint. Source may be any valid address in the same family as Flow().Source; it need not belong to LocalAddresses and is not classified as unicast, multicast, or broadcast here. It is unzoned and unmapped, and port zero is preserved on the wire. ReplyFrom has the same lifecycle, ownership, and output behavior as Reply.

type UDPForwarderResponder

type UDPForwarderResponder struct {
	// contains filtered or unexported fields
}

UDPForwarderResponder owns one detached datagram. A responder returned by Detach owns its payload snapshot; DetachForReplies omits that snapshot. It may outlive the handler because the caller, not the forwarder, owns it. It retains access to the originating forwarder's state for output, diagnostics, and Done; neither the forwarder nor the stack retains the responder. Reply calls may be repeated while active or restricted to replies; Reject and Drop are available only while active. The responder may be discarded without a terminal action.

func (*UDPForwarderResponder) Done

func (r *UDPForwarderResponder) Done() <-chan struct{}

Done is closed when the originating protocol forwarder is closed directly or by Stack.Close.

func (*UDPForwarderResponder) Drop

func (r *UDPForwarderResponder) Drop() error

Drop terminates the detached input without packet I/O. It remains valid after replies while the responder is active. It reports net.ErrClosed when the responder is already terminal or restricted to replies, including one returned by DetachForReplies.

func (*UDPForwarderResponder) Flow

Flow returns the detached datagram's original four-tuple.

func (*UDPForwarderResponder) Payload

func (r *UDPForwarderResponder) Payload() []byte

Payload returns an independently owned copy of the triggering UDP payload. It returns nil after DetachForReplies or RestrictToReplies. The caller may retain or modify a returned slice and must synchronize concurrent access.

func (*UDPForwarderResponder) Reject

func (r *UDPForwarderResponder) Reject() error

Reject terminates the detached datagram and makes a best-effort attempt to enqueue ICMP Port Unreachable without waiting for outbound capacity. Local output congestion may discard the response without error. It remains valid after replies and revalidates the forwarder and current destination policy. Once selected, the rejection decision remains terminal on output error. It reports net.ErrClosed if the responder is already terminal or restricted to replies, including one returned by DetachForReplies, or if the originating forwarder is closed.

func (*UDPForwarderResponder) Reply

func (r *UDPForwarderResponder) Reply(payload []byte) (int, error)

Reply makes one immediate best-effort output attempt for a reverse-flow datagram from Destination to Source. Use ReplyFrom to select a different source. Local output congestion may discard the datagram or any of its source fragments without error. Calls may be repeated or concurrent while the responder is active or restricted to replies, with no ordering guarantee between concurrent calls. Any call may be retried after failure; each call revalidates the forwarder and current destination policy and copies payload before returning. It uses the current Config.UDP output defaults and reports net.ErrClosed after a terminal action or when the originating forwarder is closed.

func (*UDPForwarderResponder) ReplyFrom

func (r *UDPForwarderResponder) ReplyFrom(payload []byte, source netip.AddrPort) (int, error)

ReplyFrom sends one datagram to Flow().Source with the caller-selected source IP address and UDP port. Source may be any valid address in the same family as Flow().Source; it need not belong to LocalAddresses and is not classified as unicast, multicast, or broadcast here. It is unzoned and unmapped, and port zero is preserved on the wire. Its lifecycle, ownership, concurrency, and output behavior match Reply.

func (*UDPForwarderResponder) RestrictToReplies

func (r *UDPForwarderResponder) RestrictToReplies() error

RestrictToReplies irreversibly discards the triggering payload and rejection quote while retaining Flow, Reply, ReplyFrom, and Done. On success, Reject and Drop report net.ErrClosed. Calls made after a successful restriction or on a responder returned by DetachForReplies are no-ops that return nil. It reports net.ErrClosed only if the responder is terminal and does not itself count as a terminal action. The caller must invoke it while no other responder method is running. A previously returned Payload slice remains valid and keeps its storage live for as long as the caller retains it. After this method returns, Reply and ReplyFrom may again be called concurrently.

type UDPSocketDefaults

type UDPSocketDefaults struct {
	DatagramSocketDefaults
}

UDPSocketDefaults configures policies inherited by newly created UDP sockets. Zero fields retain the package defaults.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL