design/ipam/ipam-core-library.md
This sub-design covers the core IPAM library in libcalico-go/lib/ipam/. It is the in-process API every IPAM caller goes through (CNI plugin, kube-controllers, node,
calicoctl, operator). The cross-component picture - data model, consumers, repo split - lives in the index.
The library's contract is Interface in interface.go; that file is the source of truth for method signatures and is not restated
here. A few methods carry design-relevant constraints worth calling out:
AutoAssign returns block-masked CIDRs, not /32 (or /128). Callers narrow at the boundary. This is load-bearing for the CNI plugin's per-block route programming.AssignIP enforces the target pool's allowedUses when AssignIPArgs.IntendedUse is non-empty: it fails if the pool containing the requested IP does not permit that use, mirroring the filterPoolsByUse filter AutoAssign applies. Callers that leave IntendedUse empty are exempt (back-compat). This closes a gap where a specific-IP request (e.g. the CNI ipAddrs annotation) could draw from a pool not sanctioned for its use.GetUtilization reports Capacity, InUse, Reserved and Available per pool and per block. InUse (allocated) and Reserved (covered by an IPReservation) overlap when an
address was allocated before it was reserved, so Capacity is not their sum plus Available; Available counts addresses that are neither, and is the only one of the four that
answers "how many can still be handed out". Consumers must read it rather than deriving it. Pool-level counts span the whole pool CIDR, including space no block covers yet - a
reservation over unblocked space is still unassignable - so they are computed as a set operation (pool minus reservations minus blocks, via go4.org/netipx) rather than summed
from the blocks. Reservations may overlap and nest arbitrarily, which is why a set is needed and not a sum over CIDRs.NumReservedIPsInCIDR is the pool-level reserved count on its own, for callers that already hold the IPReservations and would rather not pay for a list of every allocation
block. kube-controllers uses it for ipam_ippool_reserved from syncer-fed reservations (see ipam-gc). It takes the resources, not CIDRs, so that a variant
can take a second kind of reserving resource without reshaping its callers.ReleaseIPs takes ReleaseOptions with a sequence number; every release path must plumb it through (see CAS retry and sequence numbers).SetOwnerAttributes is KubeVirt-only and swaps owner attributes under preconditions, without releasing and re-allocating. Felix's live-migration monitor is the only non-CNI
caller.GetIPAMConfig / SetIPAMConfig read and write the v1 IPAMConfig / v3 IPAMConfiguration singleton. Field-level bounds are enforced by the CRD schema in k8s mode, but the
cross-field rules live only in SetIPAMConfig - a direct CRD write can persist a config that violates them, which the library rejects on read (see IPAMConfig).Review notes
crd.projectcalico.org/v1 types through new public APIs. The lib/v3 -> lib/internalapi rename (https://github.com/projectcalico/calico/pull/11870) exists to keep
that boundary clean.AutoAssign returning block-masked CIDRs is load-bearing for the CNI plugin's routing. Don't quietly switch to /32.GetUtilization as well as by the allocation path, or the reporting surfaces over-count free addresses. The
two must be fed from the same set of reserved CIDRs: allocation and the per-block counts share the addrFilter, and the pool-level counts use the same CIDRs as a set.reserved.go. GetUtilization and NumReservedIPsInCIDR are both thin
callers of it. Don't grow a second copy in a consumer - a reporting surface that disagrees with calicoctl ipam show is worse than no surface.Three invariants frame this section:
pending → confirmed. Pending affinities are treated as absent for ownership and routing; only confirmed affinities participate. This is
what makes the claim race resolvable.MaxBlocksPerHost caps claims, not allocations. Once a node reaches the cap it can still fill blocks it already owns. The cap is the only gate on new-block claims.StrictAffinity=true is the only way to suppress non-affine fallback. Anything that needs "borrow nothing" semantics must set it, not invent a parallel switch.AutoAssign is the hot path. Entry point is autoAssign in ipam.go. The walk splits at three chokepoints:
prepareAffinityBlocksForHost resolves the node + pools and lists existing affinities; findOrClaimBlock walks affine blocks and lazily reconstructs an IPAMBlock if an affinity
exists but the block doesn't (recovery for a node that crashed mid-claim); findUsableBlock claims a new block via the two-phase pending → confirmed protocol or, with
StrictAffinity=false, falls back to randomBlockGenerator over non-affine blocks.
A few non-obvious design points:
EmptyBlockMinReclaimAge (1 minute). The same value is what makes empty-block
release safe.min(global MaxBlocksPerHost, request-level), defaulting to 20 if both are zero. Once a node reaches it, allowNewClaim is forced false; existing blocks still
fill.Review notes
pending -> confirmed two-phase claim is intentional; don't "optimize" it away. It's what makes claim races resolvable. See
https://github.com/projectcalico/calico/pull/6003, https://github.com/projectcalico/calico/pull/1712.AffinityType; default to "host" on read. https://github.com/projectcalico/calico/pull/11179 was a crash from this assumption.ResolvePools was hand-optimized in https://github.com/projectcalico/calico/pull/9891 - preserve the fast path.MaxBlocksPerHost defaults are a recurring doc/code drift point (https://github.com/projectcalico/calico/issues/9462). If you change the default in code, update the docs in the
same PR.Two invariants frame this section:
backend.Client.Update. The retry loop is bounded; on non-conflict errors it surfaces, not swallows.All IPAMBlock, BlockAffinity, and IPAMHandle updates are CAS on resource version via backend.Client.Update. The library wraps every write in a bounded retry loop
(datastoreRetries = 100, declared in ipam.go). On ErrorResourceUpdateConflict the loop re-fetches and retries; on any other error it
exits immediately and surfaces to the caller.
Each IPAMBlock carries a SequenceNumber that updateBlock increments on every write, and per-ordinal
SequenceNumberForAllocation[ord] records the block's sequence number at allocation time. A ReleaseOptions.SequenceNumber from the caller is compared against the stored value;
on mismatch the library returns ErrorBadSequenceNumber and the IP is not released. This prevents the GC from freeing an ordinal that was reallocated to a new pod between scan
and release - the new pod's allocation bumps the block sequence number, so stale release options won't match.
Block release is parallelised per block via a semaphore sized at GOMAXPROCS.
Review notes
ReleaseOptions.SequenceNumber.*model.AllocationBlock (KVPair) before persisting. Update the value, call updateBlock, then update auxiliary state (handle, host info).
https://github.com/projectcalico/calico/pull/12697 was exactly this bug: partial mutation persisted via a retry loop after the actual write was skipped.Deallocated / Unallocated queue cycling is load-bearing for IP-reuse delay. New code that punches the bitmap directly skipping the queue breaks rate-limited reuse
(relevant to https://github.com/projectcalico/calico/issues/12638).Releasing an IP and freeing it for reuse are two distinct steps, separated by a configurable cooldown. This sits on top of the Unallocated FIFO queue: the queue cycles freed
ordinals so the longest-idle IP is reused first, and the cooldown adds a wall-clock floor on how soon any released IP can come back.
release / releaseByHandle (ipam_block.go) clear the handle association and stamp
the allocation's ReleasedAt with the current time. The ordinal stays in Allocations - the IP is no longer tied to a workload, but it is not yet available for reallocation. An
IP in this state is "in cooldown".garbageCollect deallocates IPs whose cooldown has elapsed. It moves an ordinal to Unallocated, clears its sequence number, and prunes the now-unreferenced attribute, but
only once ReleasedAt is older than IPCooldownSeconds. With IPCooldownSeconds=0 the IP is deallocated on the next GC pass. This is the only place ordinals move to
Unallocated.blockFromBackend calls garbageCollect whenever a block is loaded. On write paths (AutoAssign, release,
releaseByHandle, SetOwnerAttributes) the reclamation folds into the same CAS write, so a new allocation can reuse IPs that finished cooling down in one transaction. On
read-only paths the GC'd view is computed and then discarded - the caller sees cooled-down IPs as not-yet-reusable, but nothing is persisted.Review notes
calicoctl ipam check treats it as its own state.Unallocated FIFO, not a replacement for it. Both exist; don't remove the queue cycling thinking the timestamp covers it.A handle ID is an opaque string from the library's point of view, but its format is a convention every IPAM caller has to follow because calicoctl datastore migrate parses
the tunnel prefixes to rewrite them on node rename.
Conventions in use:
| Caller | Handle ID format |
|---|---|
| CNI workload (default) | <network-name>.<container-id> via cni-plugin/internal/pkg/utils.GetHandleID. For the default network, <network-name> is k8s-pod-network. |
| CNI workload (KubeVirt persistent) | <network-name>.<namespace>-<vm-name> so live-migrated VMs keep the same handle. |
| IPIP tunnel | ipip-tunnel-addr-<node> |
| VXLAN tunnel | vxlan-tunnel-addr-<node>; IPv6 variant is vxlan-v6-tunnel-addr-<node> (note the -v6- infix, not a suffix) |
| WireGuard tunnel | wireguard-tunnel-addr-<node>; IPv6 variant is wireguard-v6-tunnel-addr-<node> |
| Windows-reserved | literal windows-reserved-ipam-handle |
| LoadBalancer | lb-<hash>, where <hash> is the sha256 of <service>-<namespace>-<uid> (lowercased), truncated to the DNS1123 limit. Built by createHandle in kube-controllers/pkg/controllers/loadbalancer/loadbalancer_controller.go. Not to be confused with the virtual:load-balancer affinity string. |
The CNI plugin also keeps a separate workload-ID form (<namespace>.<pod>) and releases by both on DEL so that allocations made before a CRI container-ID change can still be
found - see ./ipam-cni.md.
Review notes
calicoctl ipam check classify tunnel allocations by AttributeType and Node spec respectively, not by the handle prefix - so a handle-format change doesn't touch
them, but it does need an AttributeType for the GC to recognize it.calicoctl datastore migrate (calicoctl/calicoctl/commands/datastore/migrate/migrateipam.go), which remaps tunnel-type
handles on node rename. That parser's prefix list is v4-only today (ipip-tunnel-addr-, vxlan-tunnel-addr-, wireguard-tunnel-addr-), so it already skips the *-v6-tunnel-addr-
handles - add v6 prefixes there if you touch it.IPAMConfig is a singleton CR. Defaults are applied on read by GetIPAMConfig when the CR is missing, so callers can rely on "there is always a config". Writes come from the
operator reconciling user-facing config and from end users editing the CR directly (v1 IPAMConfig or v3 IPAMConfiguration). Field-level bounds are enforced by the CRD schema in
k8s mode (kubebuilder markers on IPAMConfigurationSpec: MaxBlocksPerHost range, the KubeVirtVMAddressPersistence enum), so direct kubectl writes don't bypass those. But the
cross-field rules below live only in SetIPAMConfig, not in the CRD schema - a client that writes the CR directly can persist a config that violates them, which the library then
rejects on read:
StrictAffinity=false + AutoAllocateBlocks=false is rejected (would mean "never allocate anywhere", which is never what the user wants).MaxBlocksPerHost > 0 requires StrictAffinity=true.Fields:
| Field | Effect |
|---|---|
StrictAffinity | Disables non-affine fallback in AutoAssign. Windows forces true; see ipam-cni. |
MaxBlocksPerHost | Per-host cap on the number of affine blocks. 0 means default (20). Once a host hits the cap, allowNewClaim is forced false; existing blocks still fill. |
AutoAllocateBlocks | When false, AutoAssign will never claim a new block - only allocate from blocks the host already owns. |
KubeVirtVMAddressPersistence | Default for whether KubeVirt VM addresses survive VM restart / migration. Auto-detection is on by default. |
IPCooldownSeconds | Minimum age of a released IP before it can be reused. Release stamps the IP's ReleasedAt; garbageCollect only deallocates it once this many seconds have passed. 0 deallocates on the next GC pass. Capped at 1200. See IP release and cooldown. |
Review notes
MaxBlocksPerHost > 0 only makes sense with StrictAffinity=true; the validator enforces this. If you relax the validator, you also need to define what "borrow blocks but cap
our own" means - it currently isn't defined.SetIPAMConfig. A direct kubectl write skips the latter and
persists config the library rejects on read, so a new cross-field rule must live in the library, not on a single caller.Sentinel errors are defined in ipam_errors.go and at the top of ipam.go. The design points worth carrying here are the error
classes, not the names:
ErrorResourceUpdateConflict, errBlockClaimConflict, errStaleAffinity. The retry loop swallows these; callers never see them. Exiting the retry loop without
consuming a transient is a bug.IPAMConfigConflictError, ErrStrictAffinity, ErrNoQualifiedPool, ErrBlockLimit. Surface to callers; user-visible. New validation rules attach here.ErrorBadSequenceNumber, errBlockNotEmpty. The release path treats these as "skip this allocation and try again next sync", not as failure. A caller
that catches and proceeds anyway defeats the sequence-number / empty-block protocols.Internal sentinels (noFreeBlocksError, errBlockClaimConflict, errStaleAffinity) are not part of the public API; translate to the exported Err* values before returning.
Review notes
mustBeEmpty=true on ReleaseBlockAffinity is a hard precondition. The caller verifies emptiness; the GC's two-consecutive-empty-observations check is what gates this.ErrorBadSequenceNumber and proceeds anyway defeats the protocol. Skip and re-evaluate../ipam-datastore.md - the CAS protocol and sequence-number scheme are defined together with the backend../ipam-cni.md - the CNI plugin duplicates the handle-ID convention; changes here that affect the format need to land there too.../../libcalico-go/lib/ipam/ - a DESIGN.md stub in that directory points back here.