jasder/runc - runc - 军科开源项目托管

Commit Graph

Author	SHA1	Message	Date
Michael Crosby	e775f0fba3	Merge pull request #1526 from stevenh/logrus-v1 Updated logrus to v1	2017-07-27 13:28:55 -04:00
yangshukui	5428532bdd	remove the code that close negative descriptor Signed-off-by: yangshukui <yangshukui@huawei.com>	2017-07-24 11:10:18 +08:00
Tobias Klauser	b0d014d0e1	libcontainer: one more switch from syscall to x/sys/unix Refactor DeviceFromPath in order to get rid of package syscall and directly use the functions from x/sys/unix. This also allows to get rid of the conversion from the OS-independent file mode values (from the os package) to Linux specific values and instead let's us use the raw file mode value directly. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-07-21 16:59:15 +02:00
Steven Hartland	ee4f68e302	Updated logrus to v1 Updated logrus to use v1 which includes a breaking name change Sirupsen -> sirupsen. This includes a manual edit of the docker term package to also correct the name there too. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-07-19 15:20:56 +00:00
Daniel, Dao Quang Minh	7ab4f43a4b	Merge pull request #1519 from tklauser/moar-unix libcontainer: use additional functions and constants from x/sys/unix	2017-07-17 10:07:22 +01:00
Qiang Huang	825b5c020a	Merge pull request #1516 from cyphar/list-casting-unicode list: fix various problems with owner field	2017-07-16 14:57:20 +08:00
Tobias Klauser	4019833d46	libcontainer: use PR_SET_NO_NEW_PRIVS from x/sys/unix Use PR_SET_NO_NEW_PRIVS defined in golang.org/x/sys/unix instead of manually defining it. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-07-13 15:31:33 +02:00
Tobias Klauser	54d27bed7f	libcontainer: use ParseSocketControlMessage/ParseUnixRights from x/sys/unix Use ParseSocketControlMessage and ParseUnixRights from golang.org/x/sys/unix instead of their syscall equivalent. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-07-13 15:02:17 +02:00
Yuanhong Peng	e939079acf	Always save own namespace paths fix #1476 If containerA shares namespace, say ipc namespace, with containerB, then its ipc namespace path would be the same as containerB and be stored in `state.json`. Exec into containerA will just read the namespace paths stored in this file and join these namespaces. So, if containerB has already been stopped, `docker exec containerA` will fail. To address this issue, we should always save own namespace paths no matter if we share namespaces with other containers. Signed-off-by: Yuanhong Peng <pengyuanhong@huawei.com>	2017-07-13 16:13:05 +08:00
Michael Crosby	eb70c213ba	Update runtime-spec to rc6 Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-07-12 16:24:04 -07:00
Aleksa Sarai	7cfb107f2c	factory: use e{u,g}id as the owner of /run/runc/$id It appears as though these semantics were not fully thought out when implementing them for rootless containers. It is not necessary (and could be potentially dangerous) to set the owner of /run/ctr/$id to be the root inside the container (if user namespaces are being used). Instead, just use the e{g,u}id of runc to determine the owner. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-07-12 06:30:46 +10:00
Tobias Klauser	078e903296	libcontainer: use ioctl wrappers from x/sys/unix Use IoctlGetInt and IoctlGetTermios/IoctlSetTermios instead of manually reimplementing them. Because of unlockpt, the ioctl wrapper is still needed as it needs to pass a pointer to a value, which is not supported by any ioctl function in x/sys/unix yet. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-07-10 10:56:58 +02:00
Tobias Klauser	a380fae959	libcontainer: use Prctl() from x/sys/unix Use unix.Prctl() instead of manually reimplementing it using unix.RawSyscall. Also use unix.SECCOMP_MODE_FILTER instead of locally defining it. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-07-10 10:56:58 +02:00
Michael Crosby	5c73abbe75	Merge pull request #1450 from vrothberg/sgid-non-numeric libcontainer/user: add supplementary groups only for non-numeric users	2017-07-07 09:43:30 -07:00
Daniel, Dao Quang Minh	7139b61f7f	Merge pull request #1378 from derekwaynecarr/expose_use_hierarchy Expose memory.use_hierarchy in MemoryStats	2017-06-30 16:08:21 +01:00
Michael Crosby	fef3aced0e	Merge pull request #1460 from wking/mount-option-lazytime libcontainer/specconv/spec_linux: Add support for (no)lazytime	2017-06-29 10:06:23 -07:00
Aleksa Sarai	117c92745b	rootfs: switch ms_private remount of oldroot to ms_slave Using MS_PRIVATE meant that there was a race between the mount(2) and the umount2(2) calls where runc inadvertently has a live reference to a mountpoint that existed on the host (which the host cannot kill implicitly through an unmount and peer sharing). In particular, this means that if we have a devicemapper mountpoint and the host is trying to delete the underlying device, the delete will fail because it is "in use" during the race. While the race is _very_ small (and libdm actually retries to avoid these sorts of cases) this appears to manifest in various cases. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-06-29 01:20:23 +10:00
Justin Cormack	3d9074ead3	Update memory specs to use int64 not uint64 replace #1492 #1494 fix #1422 Since https://github.com/opencontainers/runtime-spec/pull/876 the memory specifications are now `int64`, as that better matches the visible interface where `-1` is a valid value. Otherwise finding the correct value was difficult as it was kernel dependent. Signed-off-by: Justin Cormack <justin.cormack@docker.com>	2017-06-27 12:16:07 +01:00
Justin Cormack	e1146182a8	Remove Platform as no longer in OCI spec This was never used, just validated, so was removed from spec. Signed-off-by: Justin Cormack <justin.cormack@docker.com>	2017-06-27 12:16:07 +01:00
Michael Crosby	d337d807fc	Merge pull request #1482 from tklauser/x-sys-unix-keyctl Use keyctl wrappers from x/sys/unix	2017-06-23 11:07:55 -07:00
Mrunal Patel	8e1896b3bd	Merge pull request #1491 from tklauser/unix-eventfd Use Eventfd() from golang.org/x/sys/unix	2017-06-22 19:02:44 -07:00
Michael Crosby	bd65ef625d	Merge pull request #1489 from wking/process-status libcontainer/container_linux: Consider process state (running, zombie, etc.) in runType	2017-06-21 10:24:04 -07:00
Tobias Klauser	da4cebcfe2	libcontainer: use Eventfd() from x/sys/unix Use unix.Eventfd() instead of calling manually reimplementing it using the raw syscall. Also use the correct corresponding unix.EFD_CLOEXEC flag instead of unix.FD_CLOEXEC (which can have a different value on some architectures and thus might lead to unexpected behavior). Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-06-21 10:02:00 +02:00
W. Trevor King	2bea4c897e	libcontainer/system/proc: Add Stat_t.State And Stat_t.PID and Stat_t.Name while we're at it. Then use the new .State property in runType to distinguish between running and zombie/dead processes, since kill(2) does not [1]. With this change we no longer claim Running status for zombie/dead processes. I've also removed the kill(2) call from runType. It was originally added in `13841ef3` (new-api: return the Running state only if the init process is alive, 2014-12-23), but we've been accessing /proc/[pid]/stat since `14e95b2a` (Make state detection precise, 2016-07-05, #930), and with the /stat access the kill(2) check is redundant. I also don't see much point to the previously-separate doesInitProcessExist, so I've inlined that logic in runType. It would be nice to distinguish between "/proc/[pid]/stat doesn't exist" and errors parsing its contents, but I've skipped that for the moment. The Running -> Stopped change in checkpoint_test.go is because the post-checkpoint process is a zombie, and with this commit zombie processes are Stopped (and no longer Running). [1]: https://github.com/opencontainers/runc/pull/1483#issuecomment-307527789 Signed-off-by: W. Trevor King <wking@tremily.us>	2017-06-20 16:26:55 -07:00
W. Trevor King	75d98b26b7	libcontainer: Replace GetProcessStartTime with Stat_t.StartTime And convert the various start-time properties from strings to uint64s. This removes all internal consumers of the deprecated GetProcessStartTime function. Signed-off-by: W. Trevor King <wking@tremily.us>	2017-06-20 16:26:55 -07:00
Michael Crosby	6e57120d9f	Merge pull request #1481 from elianka/dev update READ.me for new struct configs.Config.Capabilities	2017-06-20 13:15:04 -07:00
W. Trevor King	439eaa3584	libcontainer/system/proc: Add Stat and Stat_t So we can extract more than the start time with a single read. Signed-off-by: W. Trevor King <wking@tremily.us>	2017-06-14 15:28:03 -07:00
Tobias Klauser	cfe87fe3e2	Use keyctl wrappers from x/sys/unix Use KeyctlJoinSessionKeyring, KeyctlString and KeyctlSetperm from golang.org/x/sys/unix instead of manually reimplementing them. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-06-09 15:55:18 +02:00
Kang Liang	a341724c95	update READ.me for new struct configs.Config.Capabilities Signed-off-by: Kang Liang <kangliang424@gmail.com>	2017-06-09 18:47:05 +08:00
W. Trevor King	830c0d70df	libcontainer/console_linux.go: Make SaneTerminal public And use it only in local tooling that is forwarding the pseudoterminal master. That way runC no longer has an opinion on the onlcr setting for folks who are creating a terminal and detaching. They'll use --console-socket and can setup the pseudoterminal however they like without runC having an opinion. With this commit, the only cases where runC still has applies SaneTerminal is when it is the process consuming the master descriptor. Signed-off-by: W. Trevor King <wking@tremily.us>	2017-06-07 21:32:41 -07:00
Tobias Klauser	553016d7da	Use Prctl() from x/sys/unix instead of own wrapper Use unix.Prctl() instead of reimplemnting it as system.Prctl(). Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-06-07 15:03:15 +02:00
Mrunal Patel	9d6821d1b5	Merge pull request #1473 from crosbymichael/update-spec Update spec to `239c4e44f2`	2017-06-06 10:26:07 -07:00
Vladimir Stefanovic	d01050e6d4	Add support for mips/mips64 Signed-off-by: Vladimir Stefanovic <vladimir.stefanovic@imgtec.com>	2017-06-02 22:30:00 +02:00
Tobias Klauser	306b4980f7	Use NLA_* constants from x/sys/unix instead of syscall Use the NLA_ALIGNTO and NLA_HDRLEN constants from x/sys/unix instead of syscall, as the syscall package shouldn't be used anymore (except for a few exceptions). This also makes the syscall_NLA_HDRLEN workaround for gccgo unnecessary. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-06-02 10:42:11 +02:00
W. Trevor King	4f81337e95	libcontainer/specconv/spec_linux: Add support for (no)lazytime And also silent, loud, (no)iversion, and (no)acl. This is part of catching runC up with the spec, which punts valid options to mount(8) [1,2]. (no)acl is a filesystem-specific entry in mount(8), but it's represented by a MS_* flag in mount(2) so we need an entry in the translation table. [1]: https://github.com/opencontainers/runtime-spec/blame/v1.0.0-rc5/config.md#L68 [2]: https://github.com/opencontainers/runtime-spec/pull/771 Signed-off-by: W. Trevor King <wking@tremily.us>	2017-06-01 20:43:35 -07:00
Michael Crosby	18f336d23b	Merge pull request #1470 from tklauser/x-sys-unix-symlink-xattrs Use symlink xattr functions from x/sys/unix	2017-06-01 18:14:19 -07:00
Michael Crosby	854b41d81e	Update spec to `239c4e44f2` This provides updates to runc for the spec changes with *Process and OOMScoreAdj Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-06-01 16:29:47 -07:00
Tobias Klauser	d8b5c1c810	Use symlink xattr functions from x/sys/unix Use the symlink xattr syscall wrappers Lgetxattr, Llistxattr and Lsetxattr from x/sys/unix (introduced in golang/sys@b90f89a1e7) instead of providing own wrappers. Leave the functionality of system.Lgetxattr intact with respect to the retry with a larger buffer, but switch it to use unix.Lgetxattr. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-05-31 13:50:34 +02:00
Tobias Klauser	b5768387c6	Switch examples in README.md from syscall to x/sys/unix Follow commit `3d7cb4293c` ("Move libcontainer to x/sys/unix") and also move the examples in README.md from syscall to x/sys/unix. Signed-off-by: Tobias Klauser <tklauser@distanz.ch>	2017-05-30 14:50:59 +02:00
Daniel, Dao Quang Minh	67bd2ab554	Merge pull request #1442 from clnperez/libcontainer-sys-unix Move libcontainer to x/sys/unix	2017-05-26 12:18:33 +01:00
Qiang Huang	d7c264aaf1	Merge pull request #1239 from moypray/cgroup Fix setup cgroup before prestart hook	2017-05-26 09:22:49 +08:00
Michael Crosby	18cd7e06f7	Merge pull request #1372 from cloudfoundry-incubator/cpuset-mount-root Handle container creation when cgroups have already been mounted in another location	2017-05-25 09:53:57 -07:00
Christy Perez	3d7cb4293c	Move libcontainer to x/sys/unix Since syscall is outdated and broken for some architectures, use x/sys/unix instead. There are still some dependencies on the syscall package that will remain in syscall for the forseeable future: Errno Signal SysProcAttr Additionally: - os still uses syscall, so it needs to be kept for anything returning *os.ProcessState, such as process.Wait. Signed-off-by: Christy Perez <christy@linux.vnet.ibm.com>	2017-05-22 17:35:20 -05:00
Wentao Zhang	09c1f5c055	Fix setup cgroup before prestart hook * User Case: User could use prestart hook to add block devices to container. so the hook should have a way to set the permissions of the devices. Just move cgroup config operation before prestart hook will work. Signed-off-by: Wentao Zhang <zhangwentao234@huawei.com>	2017-05-19 17:53:43 +08:00
Mrunal Patel	639454475c	Merge pull request #1355 from avagin/cr-console Dump and restore containers with external terminals	2017-05-18 11:22:52 -07:00
Valentin Rothberg	77421139ab	libcontainer/user: add supplementary groups only for non-numeric users Signed-off-by: Valentin Rothberg <vrothberg@suse.com>	2017-05-16 13:54:27 +02:00
Justin Cormack	4c67360296	Clean up unix vs linux usage FreeBSD does not support cgroups or namespaces, which the code suggested, and is not supported in runc anyway right now. So clean up the file naming to use `_linux` where appropriate. Signed-off-by: Justin Cormack <justin.cormack@docker.com>	2017-05-12 17:22:09 +01:00
Qiang Huang	21ef2e3d12	Merge pull request #1410 from chchliang/statustest add createdState and runningState status testcase	2017-05-12 16:17:17 +08:00
Michael Crosby	2daa11574b	Merge pull request #1438 from hqhq/fix_rootfs_comments Fix comments about when to pivot_root	2017-05-05 20:15:49 -07:00
Qiang Huang	96e0df7633	Fix comments about when to pivot_root Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-05-06 07:59:03 +08:00
Harshal Patil	700c74cb7e	Issue #1429 : Removing check for id string length Signed-off-by: Harshal Patil <harshal.patil@in.ibm.com>	2017-05-04 09:21:29 +05:30
Harshal Patil	22953c122f	Remove redundant declaraion of namespace slice Signed-off-by: Harshal Patil <harshal.patil@in.ibm.com>	2017-05-02 10:04:57 +05:30
Andrei Vagin	73258813d3	cr: set a freezer cgroup for criu A freezer cgroup allows to dump processes faster. If a user wants to checkpoint a container and its storage, he has to pause a container, but in this case we need to pass a path to its freezer cgroup to "criu dump". Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-05-02 04:48:47 +03:00
Andrei Vagin	1c43d091a1	checkpoint: add support for containers with terminals CRIU was extended to report about orphaned master pty-s via RPC. Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-05-02 04:48:47 +03:00
Andrei Vagin	1a8b0aced5	Update criurpc Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-05-01 21:55:57 +03:00
Andrei Vagin	f8ca1926c4	libcontainer: check cpt/rst for containers with userns Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-05-01 21:45:23 +03:00
Andrei Vagin	d307e85dbb	Print a criu version in a error message Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-05-01 21:45:23 +03:00
Harshal Patil	c44d4fa6ed	Optimizing looping over namespaces Signed-off-by: Harshal Patil <harshal.patil@in.ibm.com>	2017-04-26 11:54:43 +05:30
Qiang Huang	94cfb7955b	Merge pull request #1387 from avagin/freezer Don't try to read freezer.state from the current directory	2017-04-24 20:02:45 -05:00
chchliang	4f0e6c4ef0	add createdState and runningState status testcase Signed-off-by: chchliang <chen.chuanliang@zte.com.cn>	2017-04-19 16:28:03 +08:00
Daniel, Dao Quang Minh	9f1ef73ef9	Merge pull request #1402 from chchliang/generictest add testcase in generic_error_test.go	2017-04-18 11:42:24 +01:00
chchliang	a23d7c2eab	add testcase in generic_error_test.go Signed-off-by: chchliang <chen.chuanliang@zte.com.cn>	2017-04-18 08:56:02 +08:00
Mrunal Patel	97db1eaad9	Merge pull request #1396 from harche/cstate Set container state only once during start	2017-04-17 11:32:42 -07:00
Daniel, Dao Quang Minh	13a8c5d140	Merge pull request #1365 from hqhq/use_go_selinux Use opencontainers/selinux package	2017-04-15 14:22:32 +01:00
Mrunal Patel	7814a0d14b	Merge pull request #1399 from avagin/cr-cgroup restore: apply resource limits	2017-04-13 11:28:28 -07:00
Michael Crosby	f8ce01dbdc	Merge pull request #1371 from adrianreber/master checkpoint: check if system supports pre-dumping	2017-04-12 10:08:02 -07:00
CuiHaozhi	248c586500	could load a stopped container. Signed-off-by: CuiHaozhi <cuihz@wise2c.com>	2017-04-07 07:39:41 -04:00
Andrei Vagin	57ef30a2ae	restore: apply resource limits When C/R was implemented, it was enough to call manager.Set to apply limits and to move a task. Now .Set() and .Apply() have to be called separately. Fixes: `8a740d5391` ("libcontainer: cgroups: don't Set in Apply") Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-04-07 02:47:43 +03:00
Christy Perez	fca53109c1	Fix console syscalls Fixes opencontainers/runc/issues/1364 Signed-off-by: Christy Perez <christy@linux.vnet.ibm.com>	2017-04-06 16:51:54 -05:00
Adrian Reber	273b7853c8	checkpoint: check if system supports pre-dumping Instead of relying on version numbers it is possible to check if CRIU actually supports certain features. This introduces an initial implementation to check if CRIU and the underlying kernel actually support dirty memory tracking for memory pre-dumping. Upstream CRIU also supports the lazy-page migration feature check and additional feature checks can be included in CRIU to reduce the version number parsing. There are also certain CRIU features which depend on one side on the CRIU version but also require certain kernel versions to actually work. CRIU knows if it can do certain things on the kernel it is running on and using the feature check RPC interface makes it easier for runc to decide if the criu+kernel combination will support that feature. Feature checking was introduced with CRIU 1.8. Running with older CRIU versions will ignore the feature check functionality and behave just like it used to. v2: - Do not use reflection to compare requested and responded features. Checking which feature is available is now hardcoded and needs to be adapted for every new feature check. The code is now much more readable and simpler. v3: - Move the variable criuFeat out of the linuxContainer struct, as it is not container specific. Now it is a global variable. Signed-off-by: Adrian Reber <areber@redhat.com>	2017-04-06 11:17:52 +00:00
Harshal Patil	1be5d31da2	Set container state only once during start Signed-off-by: Harshal Patil <harshal.patil@in.ibm.com>	2017-04-04 15:08:04 +05:30
Derek Carr	4d6225aec2	Expose memory.use_hierarchy in MemoryStats Signed-off-by: Derek Carr <decarr@redhat.com>	2017-03-31 13:40:34 -04:00
Aleksa Sarai	cbc4f9865a	libcontainer: rewrite cmsg to use sys/unix The original implementation is in C, which increases cognitive load and possibly might cause us problems in the future. Since sys/unix is better maintained than the syscall standard library switching makes more sense. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-30 16:03:21 +11:00
Aleksa Sarai	d04cbc49d2	rootless: add autogenerated rootless config from `runc spec` Since this is a runC-specific feature, this belongs here over in opencontainers/ocitools (which is for generic OCI runtimes). In addition, we don't create a new network namespace. This is because currently if you want to set up a veth bridge you need CAP_NET_ADMIN in both network namespaces' pinned user namespace to create the necessary interfaces in each network namespace. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:46:21 +11:00
Aleksa Sarai	76aeaf8181	libcontainer: init: fix unmapped console fchown If the stdio of the container is owned by a group which is not mapped in the user namespace, attempting to fchown the file descriptor will result in EINVAL. Counteract this by simply not doing an fchown if the group owner of the file descriptor has no host mapping according to the configured GIDMappings. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:46:21 +11:00
Aleksa Sarai	f0876b0427	libcontainer: configs: add proper HostUID and HostGID Previously Host{U,G}ID only gave you the root mapping, which isn't very useful if you are trying to do other things with the IDMaps. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:46:20 +11:00
Aleksa Sarai	baeef29858	rootless: add rootless cgroup manager The rootless cgroup manager acts as a noop for all set and apply operations. It is just used for rootless setups. Currently this is far too simple (we need to add opportunistic cgroup management), but is good enough as a first-pass at a noop cgroup manager. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:46:20 +11:00
Aleksa Sarai	d2f49696b0	runc: add support for rootless containers This enables the support for the rootless container mode. There are many restrictions on what rootless containers can do, so many different runC commands have been disabled: * runc checkpoint * runc events * runc pause * runc ps * runc restore * runc resume * runc update The following commands work: * runc create * runc delete * runc exec * runc kill * runc list * runc run * runc spec * runc state In addition, any specification options that imply joining cgroups have also been disabled. This is due to support for unprivileged subtree management not being available from Linux upstream. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:45:24 +11:00
Aleksa Sarai	6bd4bd9030	*: handle unprivileged operations and !dumpable Effectively, !dumpable makes implementing rootless containers quite hard, due to a bunch of different operations on /proc/self no longer being possible without reordering everything. !dumpable only really makes sense when you are switching between different security contexts, which is only the case when we are joining namespaces. Unfortunately this means that !dumpable will still have issues in this instance, and it should only be necessary to set !dumpable if we are not joining USER namespaces (new kernels have protections that make !dumpable no longer necessary). But that's a topic for another time. This also includes code to unset and then re-set dumpable when doing the USER namespace mappings. This should also be safe because in principle processes in a container can't see us until after we fork into the PID namespace (which happens after the user mapping). In rootless containers, it is not possible to set a non-dumpable process's /proc/self/oom_score_adj (it's owned by root and thus not writeable). Thus, it needs to be set inside nsexec before we set ourselves as non-dumpable. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-23 20:45:19 +11:00
Qiang Huang	5e7b48f7c0	Use opencontainers/selinux package It's splitted as a separate project. Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-03-23 08:21:19 +08:00
Andrei Vagin	88256d646d	Don't try to read freezer.state from the current directory If we try to pause a container on the system without freezer cgroups, we can found that runc tries to open ./freezer.state. It is obviously wrong. $ ./runc pause test no such directory for freezer.state $ echo FROZEN > freezer.state $ ./runc pause test container not running or created: paused Signed-off-by: Andrei Vagin <avagin@virtuozzo.com>	2017-03-23 01:58:45 +03:00
Daniel Dao	09c72cea69	fix panic regression when config doesnt have caps When process config doesnt specify capabilities anywhere, we should not panic because setting capabilities are optional. Signed-off-by: Daniel Dao <dqminh89@gmail.com>	2017-03-21 00:45:26 +00:00
Michael Crosby	767783a631	Merge pull request #1375 from hqhq/use_uint64_for_resources Use uint64 for resources to keep consistency with runtime-spec	2017-03-20 12:47:21 -07:00
Qiang Huang	8430cc4f48	Use uint64 for resources to keep consistency with runtime-spec Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-03-20 18:51:39 +08:00
Aleksa Sarai	c651512ad8	Revert "fix minor issue" This reverts commit `d4091ef151`. `d4091ef151` ("fix minor issue") doesn't actually make any sense, and actually makes the code more confusing. Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-03-20 12:28:43 +11:00
Qiang Huang	d270940363	Merge pull request #1356 from crosbymichael/console-socket Add separate console socket	2017-03-18 04:03:03 -05:00
Mrunal Patel	c266f1470c	Merge pull request #1373 from moypray/minor fix minor issue	2017-03-16 12:15:46 -07:00
Wentao Zhang	d4091ef151	fix minor issue When failed to attach veth pair, should remove the veth device Signed-off-by: Wentao Zhang <zhangwentao234@huawei.com>	2017-03-17 03:18:44 +08:00
Michael Crosby	957ef9cc73	Remove terminal info This maybe a nice extra but it adds complication to the usecase. The contract is listen on the socket and you get an fd to the pty master and that is that. Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-03-16 10:23:59 -07:00
Michael Crosby	00a0ecf554	Add separate console socket Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-03-16 10:23:59 -07:00
Mrunal Patel	4f903a21c4	Remove ambient build tag Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2017-03-15 11:38:43 -07:00
Mrunal Patel	4f9cb13b64	Update runtime spec to 1.0.0.rc5 Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2017-03-15 11:38:37 -07:00
Craig Furman	f5c5aac958	Create containers when cgroups already mounted Runc needs to copy certain files from the top of the cgroup cpuset hierarchy into the container's cpuset cgroup directory. Currently, runc determines which directory is the top of the hierarchy by using the parent dir of the first entry in /proc/self/mountinfo of type cgroup. This creates problems when cgroup subsystems are mounted arbitrarily in different dirs on the host. Now, we use the most deeply nested mountpoint that contains the container's cpuset cgroup directory. Signed-off-by: Konstantinos Karampogias <konstantinos.karampogias@swisscom.com> Signed-off-by: Will Martin <wmartin@pivotal.io>	2017-03-15 10:10:30 +00:00
Qiang Huang	b7932a2e07	Remove unused ExecFifoPath In container process's Init function, we use fd + execFifoFilename to open exec fifo, so this field in init config is never used. Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-03-09 10:58:16 +08:00
Qiang Huang	df4d872dd9	Merge pull request #1327 from CarltonSemple/lxd-fix Update devices_unix.go for LXD	2017-03-08 19:34:31 -06:00
Carlton-Semple	0590736890	Added comment linking to LXD issue 2825 Signed-off-by: Carlton-Semple <carlton.semple@ibm.com>	2017-03-08 10:25:37 -05:00
Qiang Huang	8773c5f9a6	Remove unused function in systemd cgroup Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-03-07 15:11:37 +08:00
Michael Crosby	49a33c41f8	Merge pull request #1344 from xuxinkun/fixCPUQuota20170224 fix cpu.cfs_quota_us changed when systemd daemon-reload using systemd.	2017-03-06 10:02:28 -08:00
xuxinkun	c44aec9b23	fix cpu.cfs_quota_us changed when systemd daemon-reload using systemd. Signed-off-by: xuxinkun <xuxinkun@gmail.com>	2017-03-06 20:08:30 +11:00
Michael Crosby	c50d024500	Merge pull request #1280 from datawolf/user user: fix the parameter error	2017-02-27 11:22:58 -08:00
Qiang Huang	fe898e7862	Fix kmem accouting when use with cgroupsPath Fixes: #1347 Fixes: #1083 The root cause of #1083 is because we're joining an existed cgroup whose kmem accouting is not initialized, and it has child cgroup or tasks in it. Fix it by checking if the cgroup is first time created, and we should enable kmem accouting if the cgroup is craeted by libcontainer with or without kmem limit configed. Otherwise we'll get issue like #1347 Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-02-25 10:58:18 -08:00
Qiang Huang	707dd48b2f	Merge pull request #1001 from x1022as/predump add pre-dump and parent-path to checkpoint	2017-02-24 10:55:06 -08:00
Aleksa Sarai	02141ce862	merge branch 'pr-1317' Closes #1317 LGTMs: @cyphar @crosbymichael	2017-02-24 08:21:58 +11:00
Qiang Huang	733563552e	Fix state when _LIBCONTAINER in environment Fixes: #1311 Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-02-22 10:35:14 -08:00
Qiang Huang	805b8c73d3	Do not create exec fifo in factory.Create It should not be binded to container creation, for example, runc restore needs to create a libcontainer.Container, but it won't need exec fifo. So create exec fifo when container is started or run, where we really need it. Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-02-22 10:34:48 -08:00
Brian Goff	d193f95d07	Don't override system error The error message added here provides no value as the caller already knows all the added details. However it is covering up the underyling system error (typically `ENOTSUP`). There is no way to handle this error before this change. Signed-off-by: Brian Goff <cpuguy83@gmail.com>	2017-02-22 09:29:38 -05:00
Michael Crosby	8438b26e9f	Merge pull request #1237 from hqhq/fix_sync_race Fix race condition when sync with child and grandchild	2017-02-20 17:16:43 -08:00
Michael Crosby	4a164a826c	Use %zu for printing of size_t values This helps fix compile warnings on some arm systems. Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-02-20 16:57:27 -08:00
Qiang Huang	a54316bae1	Fix race condition when sync with child and grandchild Fixes: #1236 Fixes: #1281 Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-02-18 20:42:08 +08:00
Qiang Huang	6b1d0e76f2	Merge pull request #1127 from boynux/fix-set-mem-to-unlimited Fixes set memory to unlimited	2017-02-16 09:51:23 +08:00
Mohammad Arab	18ebc51b3c	Reset Swap when memory is set to unlimited (-1) Kernel validation fails if memory set to -1 which is unlimited but swap is not set so. Signed-off-by: Mohammad Arab <boynux@gmail.com>	2017-02-15 08:11:57 +01:00
Carlton Semple	9a7e5a9434	Update devices_unix.go for LXD getDevices() has been updated to skip `/dev/.lxc` and `/dev/.lxd-mounts`, which was breaking privileged Docker containers running on runC, inside of LXD managed Linux Containers Signed-off-by: Carlton-Semple <carlton.semple@ibm.com>	2017-02-14 16:12:03 -05:00
Deng Guangxing	98f004182b	add pre-dump and parent-path to checkpoint CRIU gets pre-dump to complete iterative migration. pre-dump saves process memory info only. And it need parent-path to specify the former memory files. This patch add pre-dump and parent-path arguments to runc checkpoint Signed-off-by: Deng Guangxing <dengguangxing@huawei.com> Signed-off-by: Adrian Reber <areber@redhat.com>	2017-02-14 19:45:07 +08:00
Ma Shimiao	06e27471bb	support create device with type p and u Signed-off-by: Ma Shimiao <mashimiao.fnst@cn.fujitsu.com>	2017-02-10 14:45:15 +08:00
Qiang Huang	45a8341811	Small cleanup Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-02-08 15:09:06 +08:00
Qiang Huang	a8d7eb7076	Merge pull request #1314 from runcom/overlay-mounts libcontainer: rootfs_linux: support overlayfs	2017-02-08 16:17:01 +08:00
Antonio Murdaca	ca14e7b463	libcontainer: rootfs_linux: support overlayfs As the runtime-spec allows it, we want to be able to specify overlayfs mounts with: { "destination": "/etc/pki", "type": "overlay", "source": "overlay", "options": [ "lowerdir=/etc/pki:/home/amurdaca/go/src/github.com/opencontainers/runc/rootfs_fedora/etc/pki" ] }, This patch takes care of allowing overlayfs mounts. Both RO and RW should be supported. Signed-off-by: Antonio Murdaca <runcom@redhat.com>	2017-02-06 19:43:24 +01:00
Antonio Murdaca	75acc7c7c3	libcontainer: selinux: fix DupSecOpt and DisableSecOpt `label.InitLabels` takes options as a string slice in the form of: user:system_u role:system_r type:container_t level:s0:c4,c5 However, `DupSecOpt` and `DisableSecOpt` were still adding a docker specifc `label=` in front of every option. That leads to `InitLabels` not being able to correctly init selinux labels in this scenario for instance: label.InitLabels(DupSecOpt([%OPTIONS%])) if `%OPTIONS` has options prefixed with `label=`, that's going to fail. Fix this by removing that docker specific `label=` prefix. Signed-off-by: Antonio Murdaca <runcom@redhat.com>	2017-02-06 17:29:42 +01:00
Qiang Huang	7350cd8640	Merge pull request #1285 from stevenh/signal-wait Only wait for processes after delivering SIGKILL in signalAllProcesses	2017-02-06 16:41:24 +08:00
Qiang Huang	0c21b089e6	Merge pull request #1309 from stevenh/recorded-state-typo Correct docs typo for restoredState.	2017-02-04 11:51:25 +08:00
Steven Hartland	54862146c7	Correct docs typo for restoredState. Correct typo in docs for restoredState. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-02-03 16:19:01 +00:00
Steven Hartland	3f431f497e	Correct container.Destroy() docs Correct container.Destroy() docs to clarify that destroy can only operate on containers in specific states. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-02-03 16:18:29 +00:00
Qiang Huang	be33383e60	Merge pull request #1293 from stevenh/resolve-initarg Resolve InitArgs to ensure init works	2017-02-03 19:25:52 +08:00
Michael Crosby	9073486547	Merge pull request #1274 from cyphar/further-CVE-2016-9962-cleanup libcontainer: init: only pass stateDirFd when creating a container	2017-02-02 11:11:42 -08:00
Mrunal Patel	1c9c074d79	Merge pull request #1303 from runcom/revert-initlabels Revert "DupSecOpt needs to match InitLabels"	2017-02-01 10:37:16 -08:00
Steven Hartland	b9dfa444c4	Resolve InitArgs to ensure init works If a relative pathed exe is used for InitArgs init will fail to run if Cwd is not set the original path. Prevent failure of init to run by ensuring that exe in InitArgs is an absolute path. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-02-01 13:42:09 +00:00
Aleksa Sarai	e034cedce7	libcontainer: init: only pass stateDirFd when creating a container If we pass a file descriptor to the host filesystem while joining a container, there is a race condition where a process inside the container can ptrace(2) the joining process and stop it from closing its file descriptor to the stateDirFd. Then the process can access the host filesystem from that file descriptor. This was fixed in part by `5d93fed3d2` ("Set init processes as non-dumpable"), but that fix is more of a hail-mary than an actual fix for the underlying issue. To fix this, don't open or pass the stateDirFd to the init process unless we're creating a new container. A proper fix for this would be to remove the need for even passing around directory file descriptors (which are quite dangerous in the context of mount namespaces). There is still an issue with containers that have CAP_SYS_PTRACE and are using the setns(2)-style of joining a container namespace. Currently I'm not really sure how to fix it without rampant layer violation. Fixes: CVE-2016-9962 Fixes: `5d93fed3d2` ("Set init processes as non-dumpable") Signed-off-by: Aleksa Sarai <asarai@suse.de>	2017-02-02 00:41:11 +11:00
Steven Hartland	82d895fbb9	Conditionally wait for children after delivering signal When signaling children and the signal is SIGKILL wait for children otherwise conditionally wait for children which are ready to report. This reaps all children which exited due to the signal sent without blocking indefinitely. Also: * Ignore ignore ECHILD, which means the child has already gone. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-02-01 13:22:37 +00:00
Antonio Murdaca	384c1e595c	Revert "DupSecOpt needs to match InitLabels" This reverts commit `491cadac92`. Signed-off-by: Antonio Murdaca <runcom@redhat.com>	2017-02-01 09:14:20 +01:00
Mrunal Patel	510879e31f	Merge pull request #1284 from stevenh/godoc Add godoc links to README.md files	2017-01-30 10:56:58 -08:00
Daniel, Dao Quang Minh	6c22e77604	Merge pull request #1294 from stevenh/start-init-fixes Ensure pipe is always closed on error in StartInitialization	2017-01-27 16:25:44 +00:00
Qiang Huang	ed2df2906b	Merge pull request #1205 from YuPengZTE/devError fix typos by the result of golint checking	2017-01-27 21:42:18 +08:00
Mrunal Patel	c139a7c761	Merge pull request #1298 from stevenh/mention-nsenter Add nsenter details to libcontainer README.md	2017-01-25 16:25:02 -08:00
Steven Hartland	64aa78b762	Ensure pipe is always closed on error in StartInitialization Ensure that the pipe is always closed during the error processing of StartInitialization. Also: * Fix a comment typo. * Use newContainerInit directly as there's no need for i to be an initer. * Move the comment about the behaviour of Init() directly above it, clarifying what happens for all defers. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-01-25 12:36:40 +00:00
Steven Hartland	89fb8b1609	Add nsenter details to libcontainer README.md Add the import of nsenter to the example in libcontainer's README.md, as without it none of the example code works. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-01-25 01:05:36 +00:00
Justin Cormack	6ba5f5f9b8	Remove a compiler warning in some environments POSIX mandates that `cmsg_len` in `struct cmsghdr` is a `socklen_t`, which is an `unsigned int`. Musl libc as used in Alpine implements this; Glibc ignores the spec and makes it a `size_t` ie `unsigned long`. To avoid the `-Wformat=` warning from the `%lu` on Alpine, cast this to an `unsigned long` always. Signed-off-by: Justin Cormack <justin.cormack@docker.com>	2017-01-24 14:06:15 +00:00
rainrambler	4449acd306	using golang-style assignment using golang-style assignment, not the c-style Signed-off-by: Wang Anyu <wanganyu@outlook.com>	2017-01-23 14:37:16 +08:00
Steven Hartland	a887fc3f2d	Add godoc links to README.md files Add godoc links to README.md files for runc and libcontainer so its easy to access the golang documentation. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-01-21 18:21:03 +00:00
Steven Hartland	27a5447ea4	Only wait for processes after delivering SIGKILL in signalAllProcesses signalAllProcesses was making the assumption that the requested signal was SIGKILL, possibly due to the signal parameter being added at a later date, and hence it was safe to wait for all processes which is not the case. BaseContainer.Signal(s os.Signal, all bool) exposes this functionality to consumers, so an arbitrary signal could be used which is not guaranteed to make the processes exit. Correct the documentation for signalAllProcesses around the signal delivered and update it so that the wait is only performed on SIGKILL hence making it safe to process other signals without risk of blocking forever, while still maintaining compatibility to SIGKILL callers. Signed-off-by: Steven Hartland <steven.hartland@multiplay.co.uk>	2017-01-21 18:20:23 +00:00
Daniel, Dao Quang Minh	0fefa36f3a	Merge pull request #1278 from datawolf/scanner move error check out of the for loop	2017-01-20 17:49:44 +00:00
Daniel, Dao Quang Minh	b8cefd7d8f	Merge pull request #1266 from mrunalp/ignore_cgroup_v2 Ignore cgroup2 mountpoints	2017-01-20 17:26:46 +00:00
Wang Long	dde4b1a885	user: fix the parameter error The parameters passed to `GetExecUser` is not correct. Consider the following code: ``` package main import ( "fmt" "io" "os" ) func main() { passwd, err := os.Open("/etc/passwd1") if err != nil { passwd = nil } else { defer passwd.Close() } err = GetUserPasswd(passwd) if err != nil { fmt.Printf("%#v\n", err) } } func GetUserPasswd(r io.Reader) error { if r == nil { return fmt.Errorf("nil source for passwd-formatted data") } else { fmt.Printf("r = %#v\n", r) } return nil } ``` If the file `/etc/passwd1` is not exist, we expect to return `nil source for passwd-formatted data` error, and in fact, the func `GetUserPasswd` return nil. The same logic exists in runc code. this patch fix it. Signed-off-by: Wang Long <long.wanglong@huawei.com>	2017-01-19 10:02:47 +08:00
Wang Long	3a71eb0256	move error check out of the for loop The `bufio.Scanner.Scan` method returns false either by reaching the end of the input or an error. After Scan returns false, the Err method will return any error that occurred during scanning, except that if it was io.EOF, Err will return nil. We should check the error when Scan return false(out of the for loop). Signed-off-by: Wang Long <long.wanglong@huawei.com>	2017-01-18 05:02:39 +00:00
Qiang Huang	a9610f2c02	Merge pull request #1249 from datawolf/small-refactor small refactor	2017-01-13 02:04:59 -06:00
Mrunal Patel	c7ebda72ac	Add a test for testing that we ignore cgroup2 mounts Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2017-01-11 16:49:53 -08:00
Mrunal Patel	e7b57cb042	Ignore cgroup2 mountpoints Our current cgroup parsing logic assumes cgroup v1 mounts so we should ignore cgroup2 mounts for now Signed-off-by: Mrunal Patel <mrunalp@gmail.com>	2017-01-11 12:34:50 -08:00
Mrunal Patel	361bb0001a	Merge pull request #1268 from hqhq/use_source_mp Do not create cgroup dir name from combining subsystems	2017-01-11 11:34:34 -08:00
Michael Crosby	5d93fed3d2	Set init processes as non-dumpable This sets the init processes that join and setup the container's namespaces as non-dumpable before they setns to the container's pid (or any other ) namespace. This settings is automatically reset to the default after the Exec in the container so that it does not change functionality for the applications that are running inside, just our init processes. This prevents parent processes, the pid 1 of the container, to ptrace the init process before it drops caps and other sets LSMs. This patch also ensures that the stateDirFD being used is still closed prior to exec, even though it is set as O_CLOEXEC, because of the order in the kernel. https://github.com/torvalds/linux/blob/v4.9/fs/exec.c#L1290-L1318 The order during the exec syscall is that the process is set back to dumpable before O_CLOEXEC are processed. Signed-off-by: Michael Crosby <crosbymichael@gmail.com>	2017-01-11 09:56:56 -08:00
Daniel, Dao Quang Minh	2cc5a91249	Merge pull request #1260 from coolljt0725/remove_redundant Cleanup: remove redundant code	2017-01-11 17:18:15 +00:00
Qiang Huang	0599ac7d93	Do not create cgroup dir name from combining subsystems On some systems, when we mount some cgroup subsystems into a same mountpoint, the name sequence of mount options and cgroup directory name can not be the same. For example, the mount option is cpuacct,cpu, but mountpoint name is /sys/fs/cgroup/cpu,cpuacct. In current runc, we set mount destination name from combining subsystems, which comes from mount option from /proc/self/mountinfo, so in my case the name would be /sys/fs/cgroup/cpuacct,cpu, which is differernt from host, and will break some applications. Fix it by using directory name from host mountpoint. Signed-off-by: Qiang Huang <h.huangqiang@huawei.com>	2017-01-11 15:27:58 +08:00

1 2 3 4 5 ...

1046 Commits