jasder/runc - runc - 军科开源项目托管

Commit Graph

Author	SHA1	Message	Date
John Hwang	5aa0601a59	validateProcessSpec: prevent SEGV when config is valid json, but invalid. Signed-off-by: John Hwang <John.F.Hwang@gmail.com>	2020-05-18 09:38:22 -07:00
John Hwang	7fc291fd45	Replace formatted errors when unneeded Signed-off-by: John Hwang <John.F.Hwang@gmail.com>	2020-05-16 18:13:21 -07:00
Akihiro Suda	3f1e886991	Merge pull request #2391 from cyphar/devices-cgroup cgroup: devices: major cleanups and minimal transition rules	2020-05-14 09:57:06 +09:00
Kir Kolyshkin	85d4264d8a	Merge pull request #2390 from lifubang/threadedordomain cgroupv2: don't enable threaded mode by default LGTMs: AkihiroSuda, cyphar, kolyshkin	2020-05-13 14:30:25 -07:00
Kir Kolyshkin	4b71877f99	Merge pull request #2292 from Creatone/creatone/extend-intelrdt Add RDT Memory Bandwidth Monitoring (MBM) and Cache Monitoring Technology (CMT) statistics.	2020-05-13 13:33:55 -07:00
Kir Kolyshkin	41855317b6	Merge pull request #2271 from katarzyna-z/kk-cpuacct-usage-all Add reading of information from cpuacct.usage_all	2020-05-13 13:33:05 -07:00
lifubang	fe0669b26d	don't enable threaded mode by default Because in threaded mode, we can't enable the memory controller -- it isn't thread-aware. Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-13 16:27:36 +08:00
Aleksa Sarai	ba6eb28229	tests: add integration test for paused-and-updated containers Such containers should remain paused after the update. This has historically been true, but this helps ensure that the systemd cgroup changes (freezing the container during SetUnitProperties) don't break this behaviour. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:44:11 +10:00
Aleksa Sarai	4438eaa5e4	tests: add integration test for devices transition rules Unfortunately, runc update doesn't support setting devices rules directly so we have to trigger it by modifying a different rule (which happens to trigger a devices update). Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:44:11 +10:00
Aleksa Sarai	b810da1490	cgroups: systemd: make use of Device= properties It seems we missed that systemd added support for the devices cgroup, as a result systemd would actually write an allow-all rule each time you did 'runc update'* if you used the systemd cgroup driver. This is obviously ... bad and was a clear security bug. Luckily the commits which introduced this were never in an actual runc release. So we simply generate the cgroupv1-style rules (which is what systemd's DeviceAllow wants) and default to a deny-all ruleset. Unfortunately it turns out that systemd is susceptible to the same spurrious error failure that we were, so that problem is out of our hands for systemd cgroup users. However, systemd has a similar bug to the one fixed in [1]. It will happily write a disruptive deny-all rule when it is not necessary. Unfortunately, we cannot even use devices.Emulator to generate a minimal set of transition rules because the DBus API is limited (you can only clear or append to the DeviceAllow= list -- so we are forced to always clear it). To work around this, we simply freeze the container during SetUnitProperties. [1]: `afe83489d4` ("cgroupv1: devices: use minimal transition rules with devices.Emulator") Fixes: `1d4ccc8e0c` ("fix data inconsistent when runc update in systemd driven cgroup v1") Fixes: `7682a2b2a5` ("fix data inconsistent when runc update in systemd driven cgroup v2") Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:43:56 +10:00
Aleksa Sarai	afe83489d4	cgroupv1: devices: use minimal transition rules with devices.Emulator Now that all of the infrastructure for devices.Emulator is in place, we can finally implement minimal transition rules for devices cgroups. This allows for minimal disruption to running containers if a rule update is requested. Only in very rare circumstances (black-list cgroups and mode switching) will a clear-all rule be written. As a result, containers should no longer see spurious errors. A similar issue affects the cgroupv2 devices setup, but that is a topic for another time (as the solution is drastically different). Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:42:43 +10:00
Aleksa Sarai	2353ffec2b	cgroups: implement a devices cgroupv1 emulator Okay, this requires a bit of explanation. The reason for this emulation is to allow us to have seamless updates of the devices cgroup for running containers. This was triggered by several users having issues where our initial writing of a deny-all rule (in all cases) results in spurrious errors. The obvious solution would be to just remove the deny-all rule, right? Well, it turns out that runc doesn't actually control the deny-all rule because all users of runc have explicitly specified their own deny-all rule for many years. This appears to have been done to work around a bug in runc (which this series has fixed in [1]) where we would actually act as a black-list despite this being a violation of the OCI spec. This means that not adding our own deny-all rule in the case of updates won't solve the issue. However, it will also not solve the issue in several other cases (the most notable being where a container is being switched between default-permission modes). So in order to handle all of these cases, a way of tracking the relevant internal cgroup state (given a certain state of "cgroups.list" and a set of rules to apply) is necessary. That is the purpose of DevicesEmulator. Reading "devices.list" is quite important because that's the only way we can tell if it's safe to skip the troublesome deny-all rules without making potentially-dangerous assumptions about the container. We also are currently bug-compatible with the devices cgroup (namely, removing rules that don't exist or having superfluous rules all works as with the in-kernel implementation). The only exception to this is that we give an error if a user requests to revoke part of a wildcard exception, because allowing such configurations could result in security holes (cgroupv1 silently ignores such rules, meaning in white-list mode that the access is still permitted). [1]: `b2bec9806f` ("cgroup: devices: eradicate the Allow/Deny lists") Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:42:20 +10:00
Aleksa Sarai	24388be71e	configs: use different types for .Devices and .Resources.Devices Making them the same type is simply confusing, but also means that you could accidentally use one in the wrong context. This eliminates that problem. This also includes a whole bunch of cleanups for the types within DeviceRule, so that they can be used more ergonomically. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	60e21ec26e	specconv: remove default /dev/console access /dev/console is a host resouce which gives a bunch of permissions that we really shouldn't be giving to containers, not to mention that /dev/console in containers is actually /dev/pts/$n. Drop this since arguably this is a fairly scary thing to allow... Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	b2bec9806f	cgroup: devices: eradicate the Allow/Deny lists These lists have been in the codebase for a very long time, and have been unused for a large portion of that time -- specconv doesn't generate them and the only user of these flags has been tests (which doesn't inspire much confidence). In addition, we had an incorrect implementation of a white-list policy. This wasn't exploitable because all of our users explicitly specify "deny all" as the first rule, but it was a pretty glaring issue that came from the "feature" that users can select whether they prefer a white- or black- list. Fix this by always writing a deny-all rule (which is what our users were doing anyway, to work around this bug). This is one of many changes needed to clean up the devices cgroup code. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	859a780d6f	cgroups: add GetFreezerState() helper to Manager This is effectively a nicer implementation of the container.isPaused() helper, but to be used within the cgroup code for handling some fun issues we have to fix with the systemd cgroup driver. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Aleksa Sarai	a79fa7caa0	contrib: recvtty: add --no-stdin flag This is mostly just useful for testing with the "single" mode, since it allows you to run recvtty in the background without the console being closed. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2020-05-13 17:38:45 +10:00
Mrunal Patel	df3d7f673a	Merge pull request #2393 from kolyshkin/criu-pi Vagrantfile: use criu from stable repo	2020-05-12 17:48:34 -07:00
Akihiro Suda	58bf083500	Merge pull request #2400 from kolyshkin/bats-1.2.0 Dockerfile: bump bats to 1.2.0	2020-05-13 08:56:53 +09:00
Kir Kolyshkin	17aee8c432	Dockerfile: bump bats to 1.2.0 Release notes: https://github.com/bats-core/bats-core/releases/tag/v1.2.0 Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-12 11:54:17 -07:00
Akihiro Suda	2b9a36ee8c	Merge pull request #2398 from pkagrawal/master Honor spec.Process.NoNewPrivileges in specconv.CreateLibcontainerConfig	2020-05-12 15:05:55 +09:00
Mrunal Patel	867c9f5bc4	Merge pull request #2386 from kolyshkin/gordian-knot Simplify cgroup paths handling in v2 via unified v1/v2 API	2020-05-11 21:20:33 -07:00
Pradyumna Agrawal	4aa9101477	Honor spec.Process.NoNewPrivileges in specconv.CreateLibcontainerConfig The change ensures that the passed in value of NoNewPrivileges under spec.Process is reflected in the container config generated by specconv.CreateLibcontainerConfig Closes #2397 Signed-off-by: Pradyumna Agrawal <pradyumnaa@vmware.com>	2020-05-11 13:38:14 -07:00
Kir Kolyshkin	f0daf65100	Vagrantfile: use criu from stable repo CRIU 3.14 has made its way to the F32 stable repo, let's use it. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-09 13:54:27 -07:00
Kir Kolyshkin	714c91e9f7	Simplify cgroup path handing in v2 via unified API This unties the Gordian Knot of using GetPaths in cgroupv2 code. The problem is, the current code uses GetPaths for three kinds of things: 1. Get all the paths to cgroup v1 controllers to save its state (see (linuxContainer).currentState(), (LinuxFactory).loadState() methods). 2. Get all the paths to cgroup v1 controllers to have the setns process enter the proper cgroups in `(*setnsProcess).start()`. 3. Get the path to a specific controller (for example, `m.GetPaths()["devices"]`). Now, for cgroup v2 instead of a set of per-controller paths, we have only one single unified path, and a dedicated function `GetUnifiedPath()` to get it. This discrepancy between v1 and v2 cgroupManager API leads to the following problems with the code: - multiple if/else code blocks that have to treat v1 and v2 separately; - backward-compatible GetPaths() methods in v2 controllers; - - repeated writing of the PID into the same cgroup for v2; Overall, it's hard to write the right code with all this, and the code that is written is kinda hard to follow. The solution is to slightly change the API to do the 3 things outlined above in the same manner for v1 and v2: 1. Use `GetPaths()` for state saving and setns process cgroups entering. 2. Introduce and use Path(subsys string) to obtain a path to a subsystem. For v2, the argument is ignored and the unified path is returned. This commit converts all the controllers to the new API, and modifies all the users to use it. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 12:04:06 -07:00
Kir Kolyshkin	2c8d668eee	Merge pull request #2387 from kolyshkin/g-knot-prepare cgroup refactoring LGTMs: AkihiroSuda, mrunalp.	2020-05-08 12:03:22 -07:00
Kir Kolyshkin	1d143562d2	libct/cgroups/fs: access m.paths under lock 1. Prevent theoretical "concurrent map access" error to m.paths. 2. There is no need to call m.Paths -- we can access m.paths directly. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:09:55 -07:00
Kir Kolyshkin	51e1a0842d	libct/cgroups/systemd/v1: privatize v1 manager This patch was generated entirely by gorename -- nothing to review here. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:09:48 -07:00
Kir Kolyshkin	d827e323b0	libct/cgroups/systemd/v1: add NewLegacyManager Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:07:40 -07:00
Kir Kolyshkin	fc620fdf81	libct/cgroups/fs: privatize Manager and its fields This was generated entirely by gorename -- nothing to review here. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:07:00 -07:00
Kir Kolyshkin	5935bf8c21	libct/cgroups/fs: introduce NewManager() ...and use it from libcontainer/factory_linux.go. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:06:05 -07:00
Kir Kolyshkin	24f945e08d	libct/cgroups/systemd/v2: return a public interface Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:06:02 -07:00
Kir Kolyshkin	63854b0ea8	newSetnsProcess: reuse state.CgroupPaths c.cgroupManager.GetPaths() are called twice here: once in currentState() and then in newSetnsProcess(). Reuse the result of the first call, which is stored into state.CgroupPaths. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:05:59 -07:00
Kir Kolyshkin	9a3e632625	notify: simplify usage Instead of passing the whole map of paths, pass the path to the memory controller which these functions actually require. Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>	2020-05-08 10:05:58 -07:00
Aleksa Sarai	6621af89e5	merge branch 'pr-2381' Alice Frosi (3): epbf: update github.com/cilium/ebpf test: update devicefilter tests ebpf: fix big endian issue for s390x LGTMs: @AkihiroSuda @cyphar Closes #2381	2020-05-09 00:05:18 +10:00
Alice Frosi	828e4ad89d	epbf: update github.com/cilium/ebpf Update ebpf to include PR https://github.com/cilium/ebpf/pull/91. The update is needed to fix #2316. Signed-off-by: Alice Frosi <afrosi@de.ibm.com>	2020-05-08 14:20:19 +02:00
Alice Frosi	b18a9650f8	test: update devicefilter tests The test cases need to take into account the assembly modifications. The instruction: LdXMemH dst: r2 src: r1 off: 0 imm: 0 has been replaced with: LdXMemW dst: r2 src: r1 off: 0 imm: 0 And32Imm dst: r2 imm: 65535 Signed-off-by: Alice Frosi <afrosi@de.ibm.com>	2020-05-08 07:31:05 +01:00
Alice Frosi	128cb60f58	ebpf: fix big endian issue for s390x Load the full 32 bits word and take the lower 16 bits, instead of reading just 16 bits. Same fix as `07bae05e61` Signed-off-by: Alice Frosi <afrosi@de.ibm.com>	2020-05-08 07:31:05 +01:00
Kir Kolyshkin	2b31437caa	Merge pull request #2281 from AkihiroSuda/rootless-systemd cgroup v2: support rootless systemd LGTMs: kolyshkin, mrunalp	2020-05-07 21:45:52 -07:00
Mrunal Patel	47a7343182	Merge pull request #2373 from kolyshkin/logging-nits Logging nits	2020-05-07 20:54:49 -07:00
Akihiro Suda	492cfd8bf9	Merge pull request #2352 from lifubang/eventsv2 fix runc events error in cgroup v2	2020-05-08 12:51:05 +09:00
Akihiro Suda	bf15cc99b1	cgroup v2: support rootless systemd Tested with both Podman (master) and Moby (master), on Ubuntu 19.10 . $ podman --cgroup-manager=systemd run -it --rm --runtime=runc \ --cgroupns=host --memory 42m --cpus 0.42 --pids-limit 42 alpine / # cat /proc/self/cgroup 0::/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/memory.max 44040192 / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/cpu.max 42000 100000 / # cat /sys/fs/cgroup/user.slice/user-1001.slice/user@1001.service/user.slice/libpod-132ff0d72245e6f13a3bbc6cdc5376886897b60ac59eaa8dea1df7ab959cbf1c.scope/pids.max 42 Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>	2020-05-08 12:39:20 +09:00
lifubang	657407ff23	fix runc events error in cgroup v2 Signed-off-by: lifubang <lifubang@acmcoder.com>	2020-05-07 22:18:46 +08:00
Kir Kolyshkin	64416d34f3	Merge pull request #2382 from thaJeztah/bump_selinux vendor: opencontainers/selinux v1.5.1, update deprecated uses	2020-05-06 14:33:00 -07:00
Sebastiaan van Stijn	b48bbdd08d	vendor: opencontainers/selinux v1.5.1, update deprecated uses full diff: https://github.com/opencontainers/selinux/v1.4.0...v1.5.1 Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2020-05-05 15:53:40 +02:00
Katarzyna Kujawa	407e9f9d0d	Add reading of information from cpuacct.usage_all Remove logrus logs from tests Signed-off-by: Katarzyna Kujawa <katarzyna.kujawa@intel.com>	2020-05-05 08:51:12 +02:00
Mrunal Patel	a57358e016	Merge pull request #2370 from lifubang/swap0 let runc disable swap in cgroup v2	2020-05-04 16:57:12 -07:00
Kir Kolyshkin	96310f0476	Merge pull request #2377 from thaJeztah/ticks_simplify Simplify ticks, as the value is a constant	2020-05-04 15:51:31 -07:00
Sebastiaan van Stijn	402d645c5c	Simplify ticks, as the value is a constant See for example in the Musl libc source code https://git.musl-libc.org/cgit/musl/tree/src/conf/sysconf.c#n29 This removes the cgo dependency for the system package. Signed-off-by: Sebastiaan van Stijn <github@gone.nl>	2020-05-04 23:05:46 +02:00
Akihiro Suda	a0ddd02bf3	Merge pull request #2378 from thaJeztah/bump_logrus vendor: sirupsen/logrus v1.6.0	2020-05-05 02:42:16 +09:00

1 2 3 4 5 ...

4313 Commits All Branches Search

4313 Commits

All Branches