jasder/runc - runc - 军科开源项目托管

Commit Graph

Author	SHA1	Message	Date
lifubang	9087f2e827	fix path error in systemd when stopped When we use cgroup with systemd driver, the cgroup path will be auto removed by systemd when all processes exited. So we should check cgroup path exists when we access the cgroup path, for example in `kill/ps`, or else we will got an error. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-06-02 18:17:43 +08:00
Katarzyna Kujawa	92f831bf0c	Fix #2440 omit cpuacct.usage_all when not available Signed-off-by: Katarzyna Kujawa <katarzyna.kujawa@intel.com>	2020-06-02 09:24:11 +02:00
Mrunal Patel	332a84581e	Merge pull request #2443 from kolyshkin/kmem-fixup cgroupv1/systemd.Set: don't enable kernel memory acct	2020-05-31 10:04:45 -07:00
Kir Kolyshkin	3fe6e04510	cgroupv1/systemd.Set: don't enable kernel memory acct This is a regression from commit `1d4ccc8e0`. We only need to enable kernel memory accounting once, from the (legacyManager).Apply(), and there is no need to do it in (legacyManager).Set(). While at it, rename the method to better reflect what it's doing. This saves 1 call to mountinfo parser. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-29 17:54:50 -07:00
Kir Kolyshkin	3249e2379c	cgroupv1: check cpu shares in place Commit `4e65e0e90a` added a check for cpu shares. Apparently, the kernel allows to set a value higher than max or lower than min without an error, but the value read back is always within the limits. The check (which was later moved out to a separate CheckCpushares() function) is always performed after setting the cpu shares, so let's move it to the very place where it is set. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-29 16:46:28 -07:00
Kir Kolyshkin	0ac92aab3f	cgroups/fs2: make removeCgroupPath faster 1. In cases there are no sub-cgroups, a single rmdir should be faster than iterating through the list of files. 2. Use unix.Rmdir() to save one more syscall since os.Remove() tries unlink(2) first which fails on a directory, and only then tries rmdir(2). 3. Re-use rmdir. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-28 11:15:31 -07:00
Mrunal Patel	4f0bdafc8a	Merge pull request #2412 from lifubang/removecgpath remove cgroup path recursively in cgroup v2	2020-05-27 15:50:14 -07:00
Kir Kolyshkin	be5467872d	cgroupv1: minimal fix for cpu quota regression This is a quick-n-dirty fix the regression introduced by commit `06d7c1d`, which made it impossible to only set CpuQuota (without the CpuPeriod). It partially reverts the above commit, and adds a test case. The proper fix will follow. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-26 11:02:16 -07:00
lifubang	82fa194179	remove cgroup path recursively in cgroup v2 Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-26 23:35:20 +08:00
Akihiro Suda	1f737eebaa	Merge pull request #2426 from kolyshkin/mem-swap-unlim Fix some cases of swap setting	2020-05-26 14:48:59 +09:00
Akihiro Suda	7673bee6bf	Merge pull request #2395 from lifubang/updateCgroupv2 Partially revert "CreateCgroupPath: only enable needed controllers"	2020-05-25 13:56:23 +09:00
Kir Kolyshkin	3c6e8ac4d2	cgroupv2: set mem+swap to max if mem set to max ... and mem+swap is not explicitly set otherwise. This ensures compatibility with cgroupv1 controller which interprets things this way. With this fixed, we can finally enable swap tests for cgroupv2. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-22 21:32:16 -07:00
lifubang	a67dab0ac2	Revert "CreateCgroupPath: only enable needed controllers" 1. Partially revert "CreateCgroupPath: only enable needed controllers" If we update a resource which did not limited in the beginning, it will have no effective. 2. Returns err if we use an non enabled controller, or else the user may feel success, but actually there are no effective. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-21 12:17:46 +08:00
Kir Kolyshkin	d57f5bb286	cgroupv1: don't ignore MemorySwap if Memory==-1 Commit 18ebc51b3cc3 "Reset Swap when memory is set to unlimited (-1)" added handling of the case when a user updates the container limits to set memory to unlimited (-1) but do not set any other limits. Apparently, in this case, if swap limit was previously set, kernel fails to set memory.limit_in_bytes to -1 if memory.memsw.limit_in_bytes is not set to -1. What the above commit fails to handle correctly is the request when Memory is set to -1 and MemorySwap is set to some specific limit N (where N > 0). In this case, the value of N is silently discarded and MemorySwap is set to -1 instead. This is wrong thing to do, as the limit set, even if incorrectly, should not be ignored. Fix this by only assigning MemorySwap == -1 in case it was not explicitly set. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-20 17:23:40 -07:00
Kir Kolyshkin	59897367c4	cgroups/systemd: allow to set -1 as pids.limit Currently, both systemd cgroup drivers (v1 and v2) only set "TasksMax" unit property if the value > 0, so there is no way to update the limit to -1 / unlimited / infinity / max. Since systemd driver is backed by fs driver, and both fs and fs2 set the limit of -1 properly, it works, but systemd still has the old value: # runc --systemd-cgroup update $CT --pids-limit 42 # systemctl show runc-$CT.scope \| grep TasksMax TasksMax=42 # cat /sys/fs/cgroup/system.slice/runc-$CT.scope/pids.max 42 # ./runc --systemd-cgroup update $CT --pids-limit -1 # systemctl show runc-$CT.scope \| grep TasksMax= TasksMax=42 # cat /sys/fs/cgroup/system.slice/runc-xx77.scope/pids.max max Fix by changing the condition to allow -1 as a valid value. NOTE other negative values are still being ignored by systemd drivers (as it was done before). I am not sure whether this is correct, or should we return an error. A test case is added. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-20 13:20:04 -07:00
Kir Kolyshkin	06d7c1d261	systemd+cgroupv1: fix updating CPUQuotaPerSecUSec 1. do not allow to set quota without period or period without quota, as we won't be able to calculate new value for CPUQuotaPerSecUSec otherwise. 2. do not ignore setting quota to -1 when a period is not set. 3. update the test case accordingly. Note that systemd value checks will be added in the next commit. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-20 13:17:18 -07:00
Kir Kolyshkin	e4a84bea99	cgroupv2+systemd: set MemoryLow For some reason, this was not set before. Test case is added by the next commit. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-20 13:15:29 -07:00
Kir Kolyshkin	716079f95b	Merge pull request #2406 from cyphar/devices-cgroup-header cgroups: add copyright header to devices.Emulator implementation	2020-05-20 13:01:34 -07:00
Akihiro Suda	2fa3c286b5	fix "libcontainer/cgroups/fs/cpuset.go:63:14: undefined: fmt" The compilation error had ocurred because of a bad rebase during #2401 and #2413 Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>	2020-05-19 23:38:20 +09:00
Akihiro Suda	f369199ff6	Merge pull request #2413 from JFHwang/2392-spec-check Add nil check of spec.Process in validateProcessSpec()	2020-05-19 08:11:22 +09:00
Mrunal Patel	53a4649776	Merge pull request #2401 from kolyshkin/fs-cpuset-mountinfo libct/cgroup: rm GetClosestMountpointAncestor using moby/sys/mountinfo parser	2020-05-18 10:43:55 -07:00
John Hwang	7fc291fd45	Replace formatted errors when unneeded Signed-off-by: John Hwang <John.F.Hwang@gmail.com>	2020-05-16 18:13:21 -07:00
lifubang	9ad1beb40f	never write empty string to memory.swap.max Because the empty string means set swap to 0. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-16 06:52:14 +08:00
Aleksa Sarai	dc9a7879f9	cgroups: add copyright header to devices.Emulator implementation I forgot to include this in the original patchset. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-15 11:29:51 +10:00
Akihiro Suda	3f1e886991	Merge pull request #2391 from cyphar/devices-cgroup cgroup: devices: major cleanups and minimal transition rules	2020-05-14 09:57:06 +09:00
Kir Kolyshkin	2db3240f35	libct/cgroups: rm GetClosestMountpointAncestor The function GetClosestMountpointAncestor is not very efficient, does not really belong to cgroup package, and is only used once (from fs/cpuset.go). Remove it, replacing with the implementation based on moby/sys/mountinfo parser. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-13 17:32:06 -07:00
Kir Kolyshkin	f160352682	libct/cgroup: prep to rm GetClosestMountpointAncestor This function is not very efficient, does not really belong to cgroup package, and is only used once (from fs/cpuset.go). Prepare to remove it by replacing with the implementation based on the parser from github.com/moby/sys/mountinfo parser. This commit is here to make sure the proposed replacement passes the unit test. Funny, but the unit test need to be slightly modified since it supplies the wrong mountinfo (space as the first character, empty line at the end). Validated by $ go test -v -run Ance === RUN TestGetClosestMountpointAncestor --- PASS: TestGetClosestMountpointAncestor (0.00s) PASS ok github.com/opencontainers/runc/libcontainer/cgroups 0.002s Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-13 16:26:16 -07:00
Kir Kolyshkin	85d4264d8a	Merge pull request #2390 from lifubang/threadedordomain cgroupv2: don't enable threaded mode by default LGTMs: AkihiroSuda, cyphar, kolyshkin	2020-05-13 14:30:25 -07:00
Kir Kolyshkin	41855317b6	Merge pull request #2271 from katarzyna-z/kk-cpuacct-usage-all Add reading of information from cpuacct.usage_all	2020-05-13 13:33:05 -07:00
lifubang	fe0669b26d	don't enable threaded mode by default Because in threaded mode, we can't enable the memory controller -- it isn't thread-aware. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-13 16:27:36 +08:00
Aleksa Sarai	b810da1490	cgroups: systemd: make use of Device= properties It seems we missed that systemd added support for the devices cgroup, as a result systemd would actually write an allow-all rule each time you did 'runc update'* if you used the systemd cgroup driver. This is obviously ... bad and was a clear security bug. Luckily the commits which introduced this were never in an actual runc release. So we simply generate the cgroupv1-style rules (which is what systemd's DeviceAllow wants) and default to a deny-all ruleset. Unfortunately it turns out that systemd is susceptible to the same spurrious error failure that we were, so that problem is out of our hands for systemd cgroup users. However, systemd has a similar bug to the one fixed in [1]. It will happily write a disruptive deny-all rule when it is not necessary. Unfortunately, we cannot even use devices.Emulator to generate a minimal set of transition rules because the DBus API is limited (you can only clear or append to the DeviceAllow= list -- so we are forced to always clear it). To work around this, we simply freeze the container during SetUnitProperties. [1]: `afe83489d4` ("cgroupv1: devices: use minimal transition rules with devices.Emulator") Fixes: `1d4ccc8e0c` ("fix data inconsistent when runc update in systemd driven cgroup v1") Fixes: `7682a2b2a5` ("fix data inconsistent when runc update in systemd driven cgroup v2") Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:43:56 +10:00
Aleksa Sarai	afe83489d4	cgroupv1: devices: use minimal transition rules with devices.Emulator Now that all of the infrastructure for devices.Emulator is in place, we can finally implement minimal transition rules for devices cgroups. This allows for minimal disruption to running containers if a rule update is requested. Only in very rare circumstances (black-list cgroups and mode switching) will a clear-all rule be written. As a result, containers should no longer see spurious errors. A similar issue affects the cgroupv2 devices setup, but that is a topic for another time (as the solution is drastically different). Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:42:43 +10:00
Aleksa Sarai	2353ffec2b	cgroups: implement a devices cgroupv1 emulator Okay, this requires a bit of explanation. The reason for this emulation is to allow us to have seamless updates of the devices cgroup for running containers. This was triggered by several users having issues where our initial writing of a deny-all rule (in all cases) results in spurrious errors. The obvious solution would be to just remove the deny-all rule, right? Well, it turns out that runc doesn't actually control the deny-all rule because all users of runc have explicitly specified their own deny-all rule for many years. This appears to have been done to work around a bug in runc (which this series has fixed in [1]) where we would actually act as a black-list despite this being a violation of the OCI spec. This means that not adding our own deny-all rule in the case of updates won't solve the issue. However, it will also not solve the issue in several other cases (the most notable being where a container is being switched between default-permission modes). So in order to handle all of these cases, a way of tracking the relevant internal cgroup state (given a certain state of "cgroups.list" and a set of rules to apply) is necessary. That is the purpose of DevicesEmulator. Reading "devices.list" is quite important because that's the only way we can tell if it's safe to skip the troublesome deny-all rules without making potentially-dangerous assumptions about the container. We also are currently bug-compatible with the devices cgroup (namely, removing rules that don't exist or having superfluous rules all works as with the in-kernel implementation). The only exception to this is that we give an error if a user requests to revoke part of a wildcard exception, because allowing such configurations could result in security holes (cgroupv1 silently ignores such rules, meaning in white-list mode that the access is still permitted). [1]: `b2bec9806f` ("cgroup: devices: eradicate the Allow/Deny lists") Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:42:20 +10:00
Aleksa Sarai	24388be71e	configs: use different types for .Devices and .Resources.Devices Making them the same type is simply confusing, but also means that you could accidentally use one in the wrong context. This eliminates that problem. This also includes a whole bunch of cleanups for the types within DeviceRule, so that they can be used more ergonomically. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	60e21ec26e	specconv: remove default /dev/console access /dev/console is a host resouce which gives a bunch of permissions that we really shouldn't be giving to containers, not to mention that /dev/console in containers is actually /dev/pts/$n. Drop this since arguably this is a fairly scary thing to allow... Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	b2bec9806f	cgroup: devices: eradicate the Allow/Deny lists These lists have been in the codebase for a very long time, and have been unused for a large portion of that time -- specconv doesn't generate them and the only user of these flags has been tests (which doesn't inspire much confidence). In addition, we had an incorrect implementation of a white-list policy. This wasn't exploitable because all of our users explicitly specify "deny all" as the first rule, but it was a pretty glaring issue that came from the "feature" that users can select whether they prefer a white- or black- list. Fix this by always writing a deny-all rule (which is what our users were doing anyway, to work around this bug). This is one of many changes needed to clean up the devices cgroup code. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	859a780d6f	cgroups: add GetFreezerState() helper to Manager This is effectively a nicer implementation of the container.isPaused() helper, but to be used within the cgroup code for handling some fun issues we have to fix with the systemd cgroup driver. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Kir Kolyshkin	714c91e9f7	Simplify cgroup path handing in v2 via unified API This unties the Gordian Knot of using GetPaths in cgroupv2 code. The problem is, the current code uses GetPaths for three kinds of things: 1. Get all the paths to cgroup v1 controllers to save its state (see (linuxContainer).currentState(), (LinuxFactory).loadState() methods). 2. Get all the paths to cgroup v1 controllers to have the setns process enter the proper cgroups in `(*setnsProcess).start()`. 3. Get the path to a specific controller (for example, `m.GetPaths()["devices"]`). Now, for cgroup v2 instead of a set of per-controller paths, we have only one single unified path, and a dedicated function `GetUnifiedPath()` to get it. This discrepancy between v1 and v2 cgroupManager API leads to the following problems with the code: - multiple if/else code blocks that have to treat v1 and v2 separately; - backward-compatible GetPaths() methods in v2 controllers; - - repeated writing of the PID into the same cgroup for v2; Overall, it's hard to write the right code with all this, and the code that is written is kinda hard to follow. The solution is to slightly change the API to do the 3 things outlined above in the same manner for v1 and v2: 1. Use `GetPaths()` for state saving and setns process cgroups entering. 2. Introduce and use Path(subsys string) to obtain a path to a subsystem. For v2, the argument is ignored and the unified path is returned. This commit converts all the controllers to the new API, and modifies all the users to use it. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 12:04:06 -07:00
Kir Kolyshkin	2c8d668eee	Merge pull request #2387 from kolyshkin/g-knot-prepare cgroup refactoring LGTMs: AkihiroSuda, mrunalp.	2020-05-08 12:03:22 -07:00
Kir Kolyshkin	1d143562d2	libct/cgroups/fs: access m.paths under lock 1. Prevent theoretical "concurrent map access" error to m.paths. 2. There is no need to call m.Paths -- we can access m.paths directly. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:09:55 -07:00
Kir Kolyshkin	51e1a0842d	libct/cgroups/systemd/v1: privatize v1 manager This patch was generated entirely by gorename -- nothing to review here. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:09:48 -07:00
Kir Kolyshkin	d827e323b0	libct/cgroups/systemd/v1: add NewLegacyManager Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:07:40 -07:00
Kir Kolyshkin	fc620fdf81	libct/cgroups/fs: privatize Manager and its fields This was generated entirely by gorename -- nothing to review here. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:07:00 -07:00
Kir Kolyshkin	5935bf8c21	libct/cgroups/fs: introduce NewManager() ...and use it from libcontainer/factory_linux.go. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:06:05 -07:00
Kir Kolyshkin	24f945e08d	libct/cgroups/systemd/v2: return a public interface Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:06:02 -07:00
Alice Frosi	b18a9650f8	test: update devicefilter tests The test cases need to take into account the assembly modifications. The instruction: LdXMemH dst: r2 src: r1 off: 0 imm: 0 has been replaced with: LdXMemW dst: r2 src: r1 off: 0 imm: 0 And32Imm dst: r2 imm: 65535 Signed-off-by: Alice Frosi <afrosi@de.ibm.com>	2020-05-08 07:31:05 +01:00
Alice Frosi	128cb60f58	ebpf: fix big endian issue for s390x Load the full 32 bits word and take the lower 16 bits, instead of reading just 16 bits. Same fix as `07bae05e61` Signed-off-by: Alice Frosi <afrosi@de.ibm.com>	2020-05-08 07:31:05 +01:00
Akihiro Suda	bf15cc99b1	cgroup v2: support rootless systemd Tested with both Podman (master) and Moby (master), on Ubuntu 19.10 . $ podman --cgroup-manager=systemd run -it --rm --runtime=runc \ --cgroupns=host --memory 42m --cpus 0.42 --pids-limit 42 alpine / # cat /proc/self/cgroup 0::/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/memory.max 44040192 / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/cpu.max 42000 100000 / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/pids.max 42 Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>	2020-05-08 12:39:20 +09:00
Katarzyna Kujawa	407e9f9d0d	Add reading of information from cpuacct.usage_all Remove logrus logs from tests Signed-off-by: Katarzyna Kujawa <katarzyna.kujawa@intel.com>	2020-05-05 08:51:12 +02:00
Mrunal Patel	a57358e016	Merge pull request #2370 from lifubang/swap0 let runc disable swap in cgroup v2	2020-05-04 16:57:12 -07:00
Sebastiaan van Stijn	402d645c5c	Simplify ticks, as the value is a constant See for example in the Musl libc source code https://git.musl-libc.org/cgit/musl/tree/src/conf/sysconf.c#n29 This removes the cgo dependency for the system package. Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2020-05-04 23:05:46 +02:00
lifubang	a70f354680	let runc disable swap in cgroup v2 In cgroup v2, when memory and memorySwap set to the same value which is greater than zero, runc should write zero in `memory.swap.max` to disable swap. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-03 20:57:36 +08:00
Kir Kolyshkin	c3b0b13fe9	cgroups/fs2: don't always parse /proc/self/cgroup Function defaultPath always parses /proc/self/cgroup, but the resulting value is not always used. Avoid unnecessary reading/parsing by moving the code to just before its use. Modify the test case accordingly. [v2: test: use UnifiedMountpoint, skip test if not on v2] Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-28 22:16:36 -07:00
Kir Kolyshkin	0a4dcc0203	Merge pull request #2331 from lifubang/StartTransientUnit check that StartTransientUnit/StopUnit succeeds LGTMs: @AkihiroSuda @kolyshkin Closes #2313, #2309	2020-04-28 10:47:52 -07:00
lifubang	bfa1b2aab3	check that StartTransientUnit and StopUnit succeeds Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-04-28 15:46:28 +08:00
Akihiro Suda	60c647e3b8	fs2: fix cgroup.subtree_control EPERM on rootless + add CI Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>	2020-04-27 13:30:15 +09:00
Kir Kolyshkin	b19f9cecfe	Merge pull request #2343 from lifubang/updateSystemdScope fix data inconsistency when using runc update in systemd driven cgroup	2020-04-24 23:34:19 -07:00
lifubang	1d4ccc8e0c	fix data inconsistent when runc update in systemd driven cgroup v1 Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-04-23 19:32:57 +08:00
lifubang	7682a2b2a5	fix data inconsistent when runc update in systemd driven cgroup v2 Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-04-23 19:32:07 +08:00
Kir Kolyshkin	75a92ea615	cgroupv2: allow to set EnableAllDevices=true In this case we just do not install any eBPF rules checking the devices. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-22 11:05:36 -07:00
Mrunal Patel	46be7b612e	Merge pull request #2299 from kolyshkin/fs2-init-ctrl cgroupv2: fix fs2 driver initialization	2020-04-20 21:27:42 -07:00
Kir Kolyshkin	ab276b1c09	cgroups/fs2/Destroy: use Remove, ignore ENOENT 1. There is no need to try removing it recursively. 2. Do not treat ENOENT as an error (similar to fs and systemd v1 drivers). Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	4b4bc995ad	CreateCgroupPath: only enable needed controllers 1. Instead of enabling all available controllers, figure out which ones are required, and only enable those. 2. Amend all setFoo() functions to call isFooSet(). While this might seem unnecessary, it might actually help to uncover a bug. Imagine someone: - adds a cgroup.Resources.CpuFoo setting; - modifies setCpu() to apply the new setting; - but forgets to amend isCpuSet() accordingly <-- BUG In this case, a test case modifying CpuFoo will help to uncover the BUG. This is the reason why it's added. This patch could be amended by enabling controllers on a best-effort basis, i.e. : - do not return an error early if we can't enable some controllers; - if we fail to enable all controllers at once (usually because one of them can't be enabled), try enabling them one by one. Currently this is not implemented, and it's not clear whether this would be a good way to go or not. [v2: add/use is${Controller}Set() functions] [v3: document neededControllers()] [v4: drop "best-effort" part] Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	bb47e35843	cgroup/systemd: reorganize 1. Rename the files - v1.go: cgroupv1 aka legacy; - v2.go: cgroupv2 aka unified hierarchy; - unsupported.go: when systemd is not available. 2. Move the code that is common between v1 and v2 to common.go Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	de1134156b	cgroups/fs2/CreateCgroupPath: nit This slightly improves code readability. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	b5c1949f2a	cgroups/fs2/CreateCgroupPath: reinstate check This check was removed in commit `5406833a65`. Now, when this function is called from a few places, it is no longer obvious that the path always starts with /sys/fs/cgroup/, so reinstate the check just to be on the safe side. This check also ensures that elements[3:] can be used safely. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	813cb3eb94	cgroupv2: fix fs2 cgroup init fs2 cgroup driver was not working because it did not enable controllers while creating cgroup directory; instead it was merely doing MkdirAll() and gathered the list of available controllers in NewManager(). Also, cgroup should be created in Apply(), not while creating a new manager instance. To fix: 1. Move the createCgroupsv2Path function from systemd driver to fs2 driver, renaming it to CreateCgroupPath. Use in Apply() from both fs2 and systemd drivers. 2. Delay available controllers map initialization to until it is needed. With this patch: - NewManager() only performs minimal initialization (initializin m.dirPath, if not provided); - Apply() properly creates cgroup path, enabling the controllers; - m.controllers is initialized lazily on demand. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	60eaed2ed6	cgroupv2: move sanity path check to common code The fs2 cgroup driver has a sanity check for path. Since systemd driver is relying on the same path, it makes sense to move this check to the common code. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	dbeff89491	cgroupv2/systemd: privatize UnifiedManager ... and its Cgroup field. There is no sense to keep it public. This was generated by gorename. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:27:40 -07:00
Kir Kolyshkin	88c13c0713	cgroupv2: use SecureJoin in systemd driver It seems that some paths are coming from user and are therefore untrusted. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:20:22 -07:00
Kir Kolyshkin	9c80cd672d	cgroupv2: rm legacy Paths from systemd driver Having map of per-subsystem paths in systemd unified cgroups driver does not make sense and makes the code less readable. To get rid of it, move the systemd v1-or-v2 init code to libcontainer/factory_linux.go which already has a function to deduce unified path out of paths map. End result is much cleaner code. Besides, we no longer write pid to the same cgroup file 7 times in Apply() like we did before. While at it - add `rootless` flag which is passed on to fs2 manager - merge getv2Path() into GetUnifiedPath(), don't overwrite path if it is set during initialization (on Load). Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-19 16:19:51 -07:00
Kir Kolyshkin	480bca91be	cgroups/fs2: move type decl to beginning It was weird having it somewhere in the middle. No code change, just moving it around. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-18 18:43:41 -07:00
Kir Kolyshkin	353e91770b	cgroups/fs2: do not use securejoin In this very case, the code is writing to cgroup2 filesystem, and the file name is well known and can't possibly be a symlink. So, using securejoin is redundant. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-18 18:43:41 -07:00
Kir Kolyshkin	58f970a01f	cgroups/fscommon: use errors.Is This is a forgotten hunk from PR #2291. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-18 16:16:49 -07:00
Kir Kolyshkin	af6b9e7fa9	nit: do not use syscall package In many places (not all of them though) we can use `unix.` instead of `syscall.` as these are indentical. In particular, x/sys/unix defines: ```go type Signal = syscall.Signal type Errno = syscall.Errno type SysProcAttr = syscall.SysProcAttr const ENODEV = syscall.Errno(0x13) ``` and unix.Exec() calls syscall.Exec(). Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-18 16:16:49 -07:00
Akihiro Suda	9f6a2d4ddc	Merge pull request #2305 from kolyshkin/fs2-fix-default cgroupv2: fix fs2 driver default path	2020-04-16 10:16:48 +09:00
Michael Crosby	5c6216b1ed	Merge pull request #2278 from iwankgb/memory.numa_stats Exposing memory.numa_stats	2020-04-14 11:32:51 -04:00
Ted Yu	614bb96676	cgroupv2/systemd: Properly remove intermediate directory Signed-off-by: Ted Yu <yuzhihong@gmail.com>	2020-04-13 08:32:08 -07:00
Kir Kolyshkin	ea36045fe1	cgroupv2: fix fs2 driver default path When the cgroupv2 fs driver is used without setting cgroupsPath, it picks up a path from /proc/self/cgroup. On a host with systemd, such a path can look like (examples from my machines): - /user.slice/user-1000.slice/session-4.scope - /user.slice/user-1000.slice/user@1000.service/gnome-launched-xfce4-terminal.desktop-4260.scope - /user.slice/user-1000.slice/user@1000.service/gnome-terminal-server.service This cgroup already contains processes in it, which prevents to enable controllers for a sub-cgroup (writing to cgroup.subtree_control fails with EBUSY or EOPNOTSUPP). Obviously, a parent cgroup (which does not contain tasks) should be used. Fixes opencontainers/runc/issues/2298 Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-09 10:47:19 -07:00
Kenta Tada	e58a406b77	libcontainer: remove unneeded import Signed-off-by: Kenta Tada <Kenta.Tada@sony.com>	2020-04-09 20:14:39 +09:00
Michael Crosby	9a93b7378c	Merge pull request #2288 from kolyshkin/mem-swap cgroupv2: fix setting MemorySwap	2020-04-08 14:54:22 -04:00
iwankgb	7fe0a98e79	Exposing memory.numa_stats Making information on page usage by type and NUMA node available Signed-off-by: Maciej "Iwan" Iwanowski <maciej.iwanowski@intel.com>	2020-04-08 17:40:09 +02:00
Kir Kolyshkin	568cd62fa1	cgroupv2: only treat -1 as "max" Commit `6905b72154` treats all negative values as "max", citing cgroup v1 compatibility as a reason. In fact, in cgroup v1 only -1 is treated as "unlimited", and other negative values usually calse an error. Treat -1 as "max", pass other negative values as is (the error will be returned from the kernel). Fixes: `6905b72154` Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-08 04:08:49 -07:00
Kir Kolyshkin	c86be8a2c1	cgroupv2: fix setting MemorySwap The resources.MemorySwap field from OCI is memory+swap, while cgroupv2 has a separate swap limit, so subtract memory from the limit (and make sure values are set and sane). Make sure to set MemorySwapMax for systemd, too. Since systemd does not have MemorySwapMax for cgroupv1, it is only needed for v2 driver. [v2: return -1 on any negative value, add unit test] [v3: treat any negative value other than -1 as error] Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-07 20:45:53 -07:00
Giuseppe Scrivano	8b7ac5f4a5	libcontainer: use cgroups.NewStats otherwise the memoryStats and hugetlbStats maps are not initialized and GetStats() segfaults when using them. Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>	2020-04-07 09:45:57 +02:00
Mrunal Patel	0c7a9c0267	Merge pull request #2294 from tklauser/unused-consts Remove unused consts testScopeWait and testSliceWait	2020-04-06 13:26:42 -07:00
Tobias Klauser	3e678c08f9	Remove unused consts testScopeWait and testSliceWait These are unused since commit `518c855833` ("Remove libcontainer detection for systemd features") Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2020-04-03 21:09:43 +02:00
Michael Crosby	e4363b0387	Merge pull request #2291 from kolyshkin/errors-unwrap-v2 Use errors.As() and errors.Is() to unwrap errors	2020-04-03 11:46:11 -04:00
Michael Crosby	ec8c6950c7	Merge pull request #2235 from Zyqsempai/add-hugetlb-controller-to-cgroupv2 Added HugeTlb controller for cgroupv2	2020-04-03 11:15:06 -04:00
Kir Kolyshkin	b2272b2cba	libcontainer: use errors.Is() and errors.As() Make use of errors.Is() and errors.As() where appropriate to check the underlying error. The biggest motivation is to simplify the code. The feature requires go 1.13 but since merging #2256 we are already not supporting go 1.12 (which is an unsupported release anyway). Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-02 20:34:01 -07:00
Kir Kolyshkin	c39f87a47a	Revert "Merge pull request #2280 from kolyshkin/errors-unwrap" Using errors.Unwrap() is not the best thing to do, since it returns nil in case of an error which was not wrapped. More to say, errors package provides more elegant ways to check for underlying errors, such as errors.As() and errors.Is(). This reverts commit `f8e138855d`, reversing changes made to `6ca9d8e6da`. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-04-02 19:41:11 -07:00
Kir Kolyshkin	272c83e169	libct/cgroups: use errors.Unwrap Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-03-31 20:07:04 -07:00
Kir Kolyshkin	bd737f1e94	libct/cgroups/fs: use errors.Unwrap Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-03-31 20:07:04 -07:00
Kir Kolyshkin	d2dfc635ea	libct/cgroups/fs2: use errors.Unwrap Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-03-31 20:07:04 -07:00
Kir Kolyshkin	e4e35b8de8	libct/cgroups/fscommon.WriteFile: use errors.Unwrap Tested that the EINTR is still being detected: > $ go1.14 test -c # 1.14 is needed for EINTR to happen > $ sudo ./fscommon.test > INFO[0000] interrupted while writing 1063068 to /sys/fs/cgroup/memory/test-eint-89293785/memory.limit_in_bytes > PASS Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-03-31 20:07:04 -07:00
Kir Kolyshkin	66778b3c28	libct/setKernelMemory: use errors.Unwrap This simplifies code a lot. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-03-31 20:07:04 -07:00
Mrunal Patel	d05e5728aa	systemd: Lazy initialize the systemd dbus connection Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2020-03-30 15:24:06 -07:00
Mrunal Patel	33c6125da6	systemd: Export IsSystemdRunning() function Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2020-03-30 15:24:06 -07:00
Mrunal Patel	f1eea9051c	Merge pull request #2275 from kolyshkin/scan-nits bifio.Scan.Err usage nits	2020-03-27 11:41:06 -07:00
Mrunal Patel	75ff40cd28	Merge pull request #2273 from kolyshkin/v2-untangle cgroup v2 cleanups	2020-03-27 11:21:36 -07:00

1 2 3 4 5 ...

436 Commits