beenull/moby

Author	SHA1	Message	Date
Cory Snider	1931a1bdc7	libnetwork/diagnostic: lock mutex in help handler Acquire the mutex in the help handler to synchronize access to the handlers map. While a trivial issue---a panic in the request handler if the node joins a swarm at just the right time, which would only result in an HTTP 500 response---it is also a trivial race condition to fix. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-12-06 11:20:47 -05:00
Cory Snider	424ae36046	libnetwork/diagnostic: use standard http.Handler We don't need C-style callback functions which accept a void* context parameter: Go has closures. Drop the unnecessary httpHandlerCustom type and refactor the diagnostic server handler functions into closures which capture whatever context they need implicitly. If the node leaves and rejoins a swarm, the cluster agent and its associated NetworkDB are discarded and replaced with new instances. Upon rejoin, the agent registers its NetworkDB instance with the diagnostic server. These handlers would all conflict with the handlers registered by the previous NetworkDB instance. Attempting to register a second handler on a http.ServeMux with the same pattern will panic, which the diagnostic server would historically deal with by ignoring the duplicate handler registration. Consequently, the first NetworkDB instance to be registered would "stick" to the diagnostic server for the lifetime of the process, even after it is replaced with another instance. Improve duplicate-handler registration such that the most recently-registered handler for a pattern is used for all subsequent requests. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-12-06 11:19:59 -05:00
Cory Snider	757a004a90	libnetwork/diagnostic: drop Init method Fold it into the constructor, because that's what the constructor is supposed to do. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-12-04 15:13:17 -05:00
Cory Snider	f270057e0c	libnetwork/diagnostic: un-embed sync.Mutex field Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-12-04 15:13:17 -05:00
Rob Murray	964ab7158c	Explicitly set MTU on bridge devices. This is purely cosmetic - if a non-default MTU is configured, the bridge will have the default MTU=1500 until a container's 'veth' is connected and an MTU is set on the veth. That's a disconcerting, it looks like the config has been ignored - so, set the bridge's MTU explicitly. Fixes #37937 Signed-off-by: Rob Murray <rob.murray@docker.com>	2023-11-27 11:18:54 +00:00
Sebastiaan van Stijn	2f65748927	Merge pull request #46790 from corhere/libn/overlay-ipv6-vtep libnetwork/drivers/overlay: support IPv6 transport	2023-11-23 18:23:27 +01:00
Paweł Gronowski	d154421092	Merge pull request #46444 from cpuguy83/docker_info_slow Plumb context through info endpoint	2023-11-20 12:10:30 +01:00
Sebastiaan van Stijn	f13d8c2026	Merge pull request #46724 from rhansen/host_ipv6 New `host_ipv6` bridge option to SNAT IPv6 connections	2023-11-13 21:50:17 +01:00
Brian Goff	677d41aa3b	Plumb context through info endpoint I was trying to find out why `docker info` was sometimes slow so plumbing a context through to propagate trace data through. Signed-off-by: Brian Goff <cpuguy83@gmail.com>	2023-11-10 20:09:25 +00:00
Brian Goff	f0b89e63b9	Fix missing import for "scope" package I believe this happened due to conflicting PR's that got merged without CI re-running between them. Signed-off-by: Brian Goff <cpuguy83@gmail.com>	2023-11-09 22:48:01 +00:00
Brian Goff	524eef5d75	Merge pull request #46681 from corhere/libn/datastore-misc-cleanups	2023-11-09 11:31:30 -08:00
Cory Snider	33564a0c03	libnetwork/d/overlay: support IPv6 transport The forwarding database (fdb) of Linux VXLAN links are restricted to entries with destination VXLAN tunnel endpoint (VTEP) address of a single address family. Which address family is permitted is set when the link is created and cannot be modified. The overlay network driver creates VXLAN links such that the kernel only allows fdb entries to be created with IPv4 destination VTEP addresses. If the Swarm is configured with IPv6 advertise addresses, creating fdb entries for remote peers fails with EAFNOSUPPORT (address family not supported by protocol). Make overlay networks functional over IPv6 transport by configuring the VXLAN links for IPv6 VTEPs if the local node's advertise address is an IPv6 address. Make encrypted overlay networks secure over IPv6 transport by applying the iptables rules to the ip6tables when appropriate. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-11-09 12:04:47 -05:00
Cory Snider	e1d85da306	libnetwork/d/overlay: parse discovery data eagerly Parse the address strings once and use the binary representation internally. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-11-09 12:04:47 -05:00
Albin Kerouanton	d47b3ef4c9	libnet: early return from updateSvcRecord if no addr available Early return if the iface or its address is nil to make the whole function slightly easier to read. Signed-off-by: Albin Kerouanton <albinker@gmail.com>	2023-11-08 20:45:15 +01:00
Sebastiaan van Stijn	5b19725de2	Merge pull request #46668 from corhere/libn/svc-record-update-without-store libnetwork: svc record update without store	2023-11-03 13:47:12 +01:00
Cory Snider	7257c77e19	libnetwork/ipam: refactor prefix-overlap checks I am finally convinced that, given two netip.Prefix values a and b, the expression a.Contains(b.Addr()) \|\| b.Contains(a.Addr()) is functionally equivalent to a.Overlaps(b) The (netip.Prefix).Contains method works by masking the address with the prefix's mask and testing whether the remaining most-significant bits are equal to the same bits in the prefix. The (netip.Prefix).Overlaps method works by masking the longer prefix to the length of the shorter prefix and testing whether the remaining most-significant bits are equal. This is equivalent to shorterPrefix.Contains(longerPrefix.Addr()), therefore applying Contains symmetrically to two prefixes will always yield the same result as applying Overlaps to the two prefixes in either order. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-11-01 11:44:24 -04:00
Richard Hansen	808120e5b8	New `host_ipv6` bridge option to SNAT IPv6 connections Add a new `com.docker.network.host_ipv6` bridge option to compliment the existing `com.docker.network.host_ipv4` option. When set to an IPv6 address, this causes the bridge to insert `SNAT` rules instead of `MASQUERADE` rules (assuming `ip6tables` is enabled). `SNAT` makes it possible for users to control the source IP address used for outgoing connections. Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-25 20:11:49 -04:00
Richard Hansen	0cf113e250	Add unit tests for outgoing NAT rules Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-21 13:53:58 -04:00
Cory Snider	4af420f978	libnetwork/internal/kvstore: prune unused method The datastore never calls Get() due to how the cache is implemented. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:57:42 -04:00
Cory Snider	4039b9c9c4	libnetwork/datastore: drop (KVObject).DataScope() It wasn't being used for anything meaningful. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:38:39 -04:00
Cory Snider	4f4a897dda	libnetwork/datastore: drop (*Store).Scope() method It unconditionally returned scope.Local. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:38:37 -04:00
Cory Snider	4b40d82233	libnetwork/datastore: un-embed mutex from cache Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:37:12 -04:00
Cory Snider	9536fabaa8	libnetwork/datastore: minor code cleanup While there is nothing inherently wrong with goto statements, their use here is not helping with readability. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:37:12 -04:00
Cory Snider	43dccc6c1a	libnetwork/datastore: unconditionally use ds.cache ds.cache is never nil so the uncached code paths are unreachable in practice. And given how many KVObject deep-copy implementations shallow copy pointers and other reference-typed values, there is the distinct possibility that disabling the datastore cache could break things. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 12:37:10 -04:00
Cory Snider	5b3086db1f	libnetwork/datastore: prevent accidental recursion The datastore cache only uses the reference to its datastore to get a reference to the backing store. Modify the cache to take the backing store reference directly so that methods on the datastore can't get called, as that might result in infinite recursion between datastore and cache methods. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-19 11:56:08 -04:00
Cory Snider	bcca214e36	libnetwork: open-code updating svc records Inline the tortured logic for deciding when to skip updating the svc records to give us a fighting chance at deciphering the logic behind the logic and spotting logic bugs. Update the service records synchronously. The only potential for issues is if this change introduces deadlocks, which should be fixed by restrucuting the mutexes rather than papering over the issue with sketchy hacks like deferring the operation to a goroutine. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 19:51:21 -04:00
Cory Snider	33cf73f699	libnetwork: drop (*Controller).nmap Its only remaining purpose is to elide removing the endpoint from the service records if it was not previously added. Deleting the service records is an idempotent operation so it is harmless to delete service records which do not exist. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 19:46:18 -04:00
Cory Snider	804ef16822	libnetwork: only delete svc db entry on network rm The service db entry for each network is deleted by (*Controller).cleanupServiceDiscovery() when the network is deleted. There is no need to also eagerly delete it whenever the network's endpoint count drops to zero. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 19:46:18 -04:00
Cory Snider	c85398b020	libnetwork: drop vestigial endpoint-rename logic The logic to rename an endpoint includes code which would synchronize the renamed service records to peers through the distributed datastore. It would trigger the remote peers to pick up the rename by touching a datastore object which remote peers would have subscribed to events on. The code also asserts that the local peer is subscribed to updates on the network associated with the endpoint, presumably as a proxy for asserting that the remote peers would also be subscribed. https://github.com/moby/libnetwork/pull/712 Libnetwork no longer has support for distributed datastores or subscribing to datastore object updates, so this logic can be deleted. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 19:46:18 -04:00
Cory Snider	29da565133	libnetwork: change netWatch map to a set The map keys are only tested for presence. The value stored at the keys is unused. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 18:26:34 -04:00
Cory Snider	0456c0db87	libnetwork: refactor isDistributedControl() The meaning of the (*Controller).isDistributedControl() method is not immediately clear from the name, and it does not have any doc comment. It returns true if and only if the controller is neither a manager node nor an agent node -- that is, if the daemon is _not_ participating in a Swarm cluster. The method name likely comes from the old abandoned datastore-as-IPC control plane architecture for libnetwork. Refactor c.isDistributedControl() -> !c.isSwarmNode() to make it easier to understand code which consumes the method. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 17:59:19 -04:00
Cory Snider	749d4abd41	libnetwork: get rid of watchLoop goroutine Replace with roughly equivalent code which relies upon the existing mutexes for synchronization. Signed-off-by: Cory Snider <csnider@mirantis.com>	2023-10-17 17:06:52 -04:00
Richard Hansen	96f85def5b	s/HostIP/HostIPv4/ for `com.docker.network.host_ipv4` setting Rename all variables/fields/map keys associated with the `com.docker.network.host_ipv4` option from `HostIP` to `HostIPv4`. Rationale: * This makes the variable/field name consistent with the option name. * This makes the code more readable because it is clear that the variable/field does not hold an IPv6 address. This will hopefully avoid bugs like <https://github.com/moby/moby/issues/46445> in the future. * If IPv6 SNAT support is ever added, the names will be symmetric. Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 02:47:14 -04:00
Richard Hansen	2a14b6cf60	Use `iptRule` to simplify `setIcc` (code health) Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 02:47:14 -04:00
Richard Hansen	d7c6fd2f80	Move `programChainRule` logic to `iptRule` methods (code health) Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 02:47:13 -04:00
Richard Hansen	e260808a57	Move duplicate logic to `iptRule.Exists` method (code health) Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 01:41:09 -04:00
Richard Hansen	14d2535f13	Move `iptables.IPVersion` into `iptRule` struct (code health) Rather than pass an `iptables.IPVersion` value alongside every `iptRule` parameter, embed the IP version in the `iptRule` struct. Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 01:41:09 -04:00
Richard Hansen	4e219ebafb	Eliminate unnecessary `iptRule.preArgs` field (code health) That field was only used to pass `-t nat` for NAT rules. Now `-t <tableName>` (where `<tableName>` is one of the `iptables.Table` values) is always passed, eliminating the need for `preArgs`. Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 01:41:09 -04:00
Richard Hansen	4662e9889c	Simplify `setupIPTablesInternal` parameters (code health) Pass the entire `*networkConfiguration` struct to `setupIPTablesInternal` to simplify the function signature and improve code readability. Signed-off-by: Richard Hansen <rhansen@rhansen.org>	2023-10-14 01:41:09 -04:00
Bjorn Neergaard	f20abbc96c	libnetwork: use conntrack and --ctstate for all rules On modern kernels this is an alias; however newer code has preferred ctstate while older code has preferred the deprecated 'state' name. Prefer the newer name for uniformity in the rules libnetwork creates, and because some implementations/distributions of the xtables userland tools may not support the legacy alias. Signed-off-by: Bjorn Neergaard <bjorn.neergaard@docker.com>	2023-10-13 00:56:30 -06:00
Sebastiaan van Stijn	adea457841	Merge pull request #46553 from thaJeztah/no_panic libnetwork: Controller: getKeys, getPrimaryKeyTag: prevent panic and small refactor	2023-10-12 14:19:06 +02:00
Sebastiaan van Stijn	cff4f20c44	migrate to github.com/containerd/log v0.1.0 The github.com/containerd/containerd/log package was moved to a separate module, which will also be used by upcoming (patch) releases of containerd. This patch moves our own uses of the package to use the new module. Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2023-10-11 17:52:23 +02:00
Sebastiaan van Stijn	2835d1f7b2	Merge pull request #46603 from akerouanton/libnet-bridge-internal libnet/d/bridge: Don't set container's gateway when network is internal	2023-10-11 17:07:02 +02:00
Sebastiaan van Stijn	26c5d1ea0d	Merge pull request #46551 from akerouanton/libnet-resolver-otel libnet: add OTEL tracing to the embedded DNS	2023-10-11 17:03:30 +02:00
Albin Kerouanton	37ca57e9d5	libnet/d/bridge: inline error checks Signed-off-by: Albin Kerouanton <albinker@gmail.com>	2023-10-10 10:46:44 +02:00
Albin Kerouanton	cbc2a71c27	libnet/d/bridge: Don't set container's gateway when network is internal So far, internal networks were only isolated from the host by iptables DROP rules. As a consequence, outbound connections from containers would timeout instead of being "rejected" through an immediate ICMP dest/port unreachable, a TCP RST or a failing `connect` syscall. This was visible when internal containers were trying to resolve a domain that don't match any container on the same network (be it a truly "external" domain, or a container that don't exist/is dead). In that case, the embedded resolver would try to forward DNS queries for the different values of resolv.conf `search` option, making DNS resolution slow to return an error, and the slowness being exacerbated by some libc implementations. This change makes `connect` syscall to return ENETUNREACH, and thus solves the broader issue of failing fast when external connections are attempted. Signed-off-by: Albin Kerouanton <albinker@gmail.com>	2023-10-09 13:57:54 +02:00
Albin Kerouanton	2c4551d86d	libnet: resolver: remove direct use of logrus This causes logs written through `r.log(ctx)` to not end in OTEL traces. Signed-off-by: Albin Kerouanton <albinker@gmail.com>	2023-10-06 19:14:48 +02:00
Albin Kerouanton	4de8459265	libnet: add OTEL tracing to the embedded DNS This change creates a few OTEL spans and plumb context through the DNS resolver and DNS backends (ie. Sandbox and Network). This should help better understand how much lock contention impacts performance, and help debug issues related to DNS queries (we basically have no visibility into what's happening here right now). Signed-off-by: Albin Kerouanton <albinker@gmail.com>	2023-10-06 19:14:48 +02:00
Sebastiaan van Stijn	dcc75e1563	libnetwork: Controller: agentInit, agentDriverNotify rm intermediate vars Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2023-09-27 12:08:28 +02:00
Sebastiaan van Stijn	a384102fdf	libnetwork/datastore: Store.Map, Store.List: remove intermediate vars Inline the closures, and rename a var to be more descriptive. Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2023-09-27 12:07:31 +02:00

1 2 3 4 5 ...

3252 commits