gvisor - Container Runtime Sandbox

Age	Commit message (Collapse)	Author
2021-03-24	Add POLLRDNORM/POLLWRNORM support.	Bhasker Hariharan
	On Linux these are meant to be equivalent to POLLIN/POLLOUT. Rather than hack these on in sys_poll etc it felt cleaner to just cleanup the call sites to notify for both events. This is what linux does as well. Fixes #5544 PiperOrigin-RevId: 364859977
2021-03-23	Merge pull request #5677 from avagin:kvm-mmio	gVisor bot
	PiperOrigin-RevId: 364728696
2021-03-23	Move the code that manages floating-point state to a separate package	Andrei Vagin
	This change is inspired by Adin's cl/355256448. PiperOrigin-RevId: 364695931
2021-03-23	setgid directory support in goferfs	Kevin Krakauer
	Also adds support for clearing the setuid bit when appropriate (writing, truncating, changing size, changing UID, or changing GID). VFS2 only. PiperOrigin-RevId: 364661835
2021-03-22	Avoid calling sync on each write in writethrough mode.	Nicolas Lacasse
	PiperOrigin-RevId: 364370595
2021-03-18	Translate syserror when validating partial IO errors	Fabricio Voznika
	syserror allows packages to register translators for errors. These translators should be called prior to checking if the error is valid, otherwise it may not account for possible errors that can be returned from different packages, e.g. safecopy.BusError => syserror.EFAULT. Second attempt, it passes tests now :-) PiperOrigin-RevId: 363714508
2021-03-16	kvm: prefault a floating point state before restoring it	Andrei Vagin
	If physical pages of a memory region are not mapped yet, the kernel will trigger KVM_EXIT_MMIO and we will map physical pages in bluepillHandler(). An instruction that triggered a fault will not be re-executed, it will be emulated in the kernel, but it can't emulate complex instructions like xsave, xrstor. We can touch the memory with simple instructions to workaround this problem.
2021-03-16	setgid directory support in overlayfs	Kevin Krakauer
	PiperOrigin-RevId: 363276495
2021-03-15	Turn sys_thread constants into variables.	Etienne Perot
	PiperOrigin-RevId: 363092268
2021-03-15	Make netstack (//pkg/tcpip) buildable for 32 bit	Kevin Krakauer
	Doing so involved breaking dependencies between //pkg/tcpip and the rest of gVisor, which are discouraged anyways. Tested on the Go branch via: gvisor.dev/gvisor/pkg/tcpip/... Addresses #1446. PiperOrigin-RevId: 363081778
2021-03-15	[op] Make gofer client handle return partial write length when err is nil.	Ayush Ranjan
	If there was a partial write (when not using the host FD) which did not generate an error, we were incorrectly returning the number of bytes attempted to write instead of the number of bytes actually written. PiperOrigin-RevId: 363058989
2021-03-15	Merge pull request #5618 from iangudger:unix-transport-race	gVisor bot
	PiperOrigin-RevId: 362999220
2021-03-11	fusefs: Implement default_permissions and allow_other mount options.	Rahat Mahmood
	By default, fusefs defers node permission checks to the server. The default_permissions mount option enables the usual unix permission checks based on the node owner and mode bits. Previously fusefs was incorrectly checking permissions unconditionally. Additionally, fusefs should restrict filesystem access to processes started by the mount owner to prevent the fuse daemon from gaining priviledge over other processes. The allow_other mount option overrides this behaviour. Previously fusefs was incorrectly skipping this check. Updates #3229 PiperOrigin-RevId: 362419092
2021-03-11	Clear Merkle tree files in RuntimeEnable mode	Chong Cai
	The Merkle tree files need to be cleared before enabling to avoid redundant content. PiperOrigin-RevId: 362409591
2021-03-11	Report filesystem-specific mount options.	Rahat Mahmood
	PiperOrigin-RevId: 362406813
2021-03-08	Implement /proc/sys/net/ipv4/ip_local_port_range	Kevin Krakauer
	Speeds up the socket stress tests by a couple orders of magnitude. PiperOrigin-RevId: 361721050
2021-03-05	Implement IterDirent in verity fs	Chong Cai
	PiperOrigin-RevId: 361196154
2021-03-04	Fix race in unix socket transport.	Ian Gudger
	transport.baseEndpoint.receiver and transport.baseEndpoint.connected are protected by transport.baseEndpoint.Mutex. In order to access them without holding the mutex, we must make a copy. Notifications must be sent without holding the mutex, so we need the values without holding the mutex.
2021-03-03	Add checklocks analyzer.	Bhasker Hariharan
	This validates that struct fields if annotated with "// checklocks:mu" where "mu" is a mutex field in the same struct then access to the field is only done with "mu" locked. All types that are guarded by a mutex must be annotated with // +checklocks:<mutex field name> For more details please refer to README.md. PiperOrigin-RevId: 360729328
2021-03-03	Export stats that were forgotten	Arthur Sfez
	While I'm here, simplify the comments and unify naming of certain stats across protocols. PiperOrigin-RevId: 360728849
2021-03-03	[op] Replace syscall package usage with golang.org/x/sys/unix in pkg/.	Ayush Ranjan
	The syscall package has been deprecated in favor of golang.org/x/sys. Note that syscall is still used in the following places: - pkg/sentry/socket/hostinet/stack.go: some netlink related functionalities are not yet available in golang.org/x/sys. - syscall.Stat_t is still used in some places because os.FileInfo.Sys() still returns it and not unix.Stat_t. Updates #214 PiperOrigin-RevId: 360701387
2021-02-25	Implement SEM_STAT_ANY cmd of semctl.	Jing Chen
	PiperOrigin-RevId: 359591577
2021-02-24	Kernfs should not try to rename a file to itself.	Nicolas Lacasse
	One precondition of VFS.PrepareRenameAt is that the `from` and `to` dentries are not the same. Kernfs was not checking this, which could lead to a deadlock. PiperOrigin-RevId: 359385974
2021-02-24	Use mapped device number + topmost inode number for all files in VFS2 overlay.	Jamie Liu
	Before this CL, VFS2's overlayfs uses a single private device number and an autoincrementing generated inode number for directories; this is consistent with Linux's overlayfs in the non-samefs non-xino case. However, this breaks some applications more consistently than on Linux due to more aggressive caching of Linux overlayfs dentries. Switch from using mapped device numbers + the topmost layer's inode number for just non-copied-up non-directory files, to doing so for all files. This still allows directory dev/ino numbers to change across copy-up, but otherwise keeps them consistent. Fixes #5545: ``` $ docker run --runtime=runsc-vfs2-overlay --rm ubuntu:focal bash -c "mkdir -p 1/2/3/4/5/6/7/8 && rm -rf 1 && echo done" done ``` PiperOrigin-RevId: 359350716
2021-02-24	Merge pull request #5519 from dqminh:runsc-ps-pids	gVisor bot
	PiperOrigin-RevId: 359334029
2021-02-24	return root pids with runsc ps	Daniel Dao
	`runsc ps` currently return pid for a task's immediate pid namespace, which is confusing when there're multiple pid namespaces. We should return only pids in the root namespace. Before: ``` 1000 1 0 0 ? 02:24 250ms chrome 1000 1 0 0 ? 02:24 40ms dumb-init 1000 1 0 0 ? 02:24 240ms chrome 1000 2 1 0 ? 02:24 2.78s node ``` After: ``` UID PID PPID C TTY STIME TIME CMD 1000 1 0 0 ? 12:35 0s dumb-init 1000 2 1 7 ? 12:35 240ms node 1000 13 2 21 ? 12:35 2.33s chrome 1000 27 13 3 ? 12:35 260ms chrome ``` Signed-off-by: Daniel Dao <dqminh@cloudflare.com>
2021-02-24	Add YAMA security module restrictions on ptrace(2).	Dean Deng
	Restrict ptrace(2) according to the default configurations of the YAMA security module (mode 1), which is a common default among various Linux distributions. The new access checks only permit the tracer to proceed if one of the following conditions is met: a) The tracer is already attached to the tracee. b) The target is a descendant of the tracer. c) The target has explicitly given permission to the tracer through the PR_SET_PTRACER prctl. d) The tracer has CAP_SYS_PTRACE. See security/yama/yama_lsm.c for more details. Note that these checks are added to CanTrace, which is checked for PTRACE_ATTACH as well as some other operations, e.g., checking a process' memory layout through /proc/[pid]/mem. Since this patch adds restrictions to ptrace, it may break compatibility for applications run by non-root users that, for instance, rely on being able to trace processes that are not descended from the tracer (e.g., `gdb -p`). YAMA restrictions can be turned off by setting /proc/sys/kernel/yama/ptrace_scope to 0, or exceptions can be made on a per-process basis with the PR_SET_PTRACER prctl. Reported-by: syzbot+622822d8bca08c99e8c8@syzkaller.appspotmail.com PiperOrigin-RevId: 359237723
2021-02-24	Use async task context for async IO.	Dean Deng
	PiperOrigin-RevId: 359235699
2021-02-22	unix: sendmmsg and recvmsg have to cap a number of message to UIO_MAXIOV	Andrei Vagin
	Reported-by: syzbot+f2489ba0b999a45d1ad1@syzkaller.appspotmail.com PiperOrigin-RevId: 358866218
2021-02-19	Don't hold baseEndpoint.mu while calling EventUpdate().	Nicolas Lacasse
	This removes a three-lock deadlock between fdnotifier.notifier.mu, epoll.EventPoll.listsMu, and baseEndpoint.mu. A lock order comment was added to epoll/epoll.go. Also fix unsafe access of baseEndpoint.connected/receiver. PiperOrigin-RevId: 358515191
2021-02-19	control.Proc.Exec should default to root pid namespace if none provided.	Nicolas Lacasse
	PiperOrigin-RevId: 358445320
2021-02-18	Make socketops reflect correct sndbuf value for host UDS.	Bhasker Hariharan
	Also skips a test if the setsockopt to increase send buffer did not result in an increase. This is possible when the underlying socket is a host backed unix domain socket as in such cases gVisor does not permit increasing SO_SNDBUF. PiperOrigin-RevId: 358285158
2021-02-18	Bump build constraints to Go 1.18	Michael Pratt
	These are bumped to allow early testing of Go 1.17. Use will be audited closer to the 1.17 release. PiperOrigin-RevId: 358278615
2021-02-18	Validate IGMP packets	Arthur Sfez
	This change also adds support for Router Alert option processing on incoming packets, a new stat for Router Alert option, and exports all the IP-option related stats. Fixes #5491 PiperOrigin-RevId: 358238123
2021-02-17	Move Name() out of netstack Matcher. It can live in the sentry.	Kevin Krakauer
	PiperOrigin-RevId: 358078157
2021-02-17	Add gohacks.Slice/StringHeader.	Jamie Liu
	See https://github.com/golang/go/issues/19367 for rationale. Note that the upstream decision arrived at in that thread, while useful for some of our use cases, doesn't account for all of our SliceHeader use cases (we often use SliceHeader to extract pointers from slices in a way that avoids bounds checking and/or handles nil slices correctly) and also doesn't exist yet. PiperOrigin-RevId: 358071574
2021-02-17	Check for directory emptiness in VFS1 overlay rmdir().	Jamie Liu
	Note that this CL reorders overlayEntry.copyMu before overlayEntry.dirCacheMu in the overlayFileOperations.IterateDir() => readdirEntries() path - but this lock ordering is already required by overlayRemove/Bind() => overlayEntry.markDirectoryDirty(), so this actually just fixes an inconsistency. PiperOrigin-RevId: 358047121
2021-02-11	[rack] TLP: ACK Processing and PTO scheduling.	Ayush Ranjan
	This change implements TLP details enumerated in https://tools.ietf.org/html/draft-ietf-tcpm-rack-08#section-7.5.3 Fixes #5085 PiperOrigin-RevId: 357125037
2021-02-11	Unconditionally check for directory-ness in overlay.filesystem.UnlinkAt().	Jamie Liu
	PiperOrigin-RevId: 357106080
2021-02-11	Internal change.	gVisor bot
	PiperOrigin-RevId: 357090170
2021-02-11	Implement semtimedop.	Jing Chen
	PiperOrigin-RevId: 357031904
2021-02-11	Assign controlling terminal when tty is opened and support NOCTTY	Kevin Krakauer
	PiperOrigin-RevId: 357015186
2021-02-10	Support setgid directories in tmpfs and kernfs	Kevin Krakauer
	PiperOrigin-RevId: 356868412
2021-02-10	Don't allow to umount the namespace root mount	Andrei Vagin
	Linux does the same thing. Reported-by: syzbot+6c79385c930c929d1d9e@syzkaller.appspotmail.com PiperOrigin-RevId: 356854562
2021-02-10	Merge pull request #5267 from lubinszARM:pr_usr_lazy_fp	gVisor bot
	PiperOrigin-RevId: 356762859
2021-02-09	Add support for setting SO_SNDBUF for unix domain sockets.	Bhasker Hariharan
	The limits for snd/rcv buffers for unix domain socket is controlled by the following sysctls on linux - net.core.rmem_default - net.core.rmem_max - net.core.wmem_default - net.core.wmem_max Today in gVisor we do not expose these sysctls but we do support setting the equivalent in netstack via stack.Options() method. But AF_UNIX sockets in gVisor can be used without netstack, with hostinet or even without any networking stack at all. Which means ideally these sysctls need to live as globals in gVisor. But rather than make this a big change for now we hardcode the limits in the AF_UNIX implementation itself (which in itself is better than where we were before) where it SO_SNDBUF was hardcoded to 16KiB. Further we bump the initial limit to a default value of 208 KiB to match linux from the paltry 16 KiB we use today. Updates #5132 PiperOrigin-RevId: 356665498
2021-02-09	Add cleanup TODO for integer-based proc files.	Dean Deng
	PiperOrigin-RevId: 356645022
2021-02-09	Collapse code that always returns error	Tamir Duberstein
	PiperOrigin-RevId: 356536548
2021-02-09	kernel: reparentLocked has to update children maps of old and new parents	Andrei Vagin
	Reported-by: syzbot+9ffc71246fe72c73fc25@syzkaller.appspotmail.com PiperOrigin-RevId: 356536113
2021-02-09	pipe: writeLocked has to return ErrWouldBlock if the pipe is full	Andrei Vagin
	PiperOrigin-RevId: 356450303