gvisor - Container Runtime Sandbox

Age	Commit message (Collapse)	Author
2020-06-19	Implement UDP cheksum verification.	gVisor bot
	Test: - TestIncrementChecksumErrors Fixes #2943 PiperOrigin-RevId: 317348158
2020-06-19	Fix vfs2 handling of preadv2/pwritev2 flags.	Dean Deng
	Check for unsupported flags, and silently support RWF_HIPRI by doing nothing. From pkg/abi/linux/file.go: "gVisor does not implement the RWF_HIPRI feature, but the flag is accepted as a valid flag argument for preadv2/pwritev2." Updates #2923. PiperOrigin-RevId: 317330631
2020-06-19	Don't adjust parent link count if we replace a child dir with another.	Dean Deng
	Updates #2923. PiperOrigin-RevId: 317314460
2020-06-19	Support all seek options in gofer specialFileFD.Seek.	Dean Deng
	Updates #2923. PiperOrigin-RevId: 317298186
2020-06-19	Fix synthetic file bugs in gofer fs.	Dean Deng
	Always check if a synthetic file already exists at a location before creating a file there, and do not try to delete synthetic gofer files from the remote fs. This fixes runsc_ptrace socket tests that create/unlink synthetic, named socket files. Updates #2923. PiperOrigin-RevId: 317293648
2020-06-18	Fix vfs2 tmpfs link permission checks.	Dean Deng
	Updates #2923. PiperOrigin-RevId: 317246916
2020-06-18	socket/unix: (*connectionedEndpoint).State() has to take the endpoint lock	Andrei Vagin
	It accesses e.receiver which is protected by the endpoint lock. WARNING: DATA RACE Write at 0x00c0006aa2b8 by goroutine 189: pkg/sentry/socket/unix/transport.(connectionedEndpoint).Connect.func1() pkg/sentry/socket/unix/transport/connectioned.go:359 +0x50 pkg/sentry/socket/unix/transport.(connectionedEndpoint).BidirectionalConnect() pkg/sentry/socket/unix/transport/connectioned.go:327 +0xa3c pkg/sentry/socket/unix/transport.(connectionedEndpoint).Connect() pkg/sentry/socket/unix/transport/connectioned.go:363 +0xca pkg/sentry/socket/unix.(socketOpsCommon).Connect() pkg/sentry/socket/unix/unix.go:420 +0x13a pkg/sentry/socket/unix.(SocketOperations).Connect() <autogenerated>:1 +0x78 pkg/sentry/syscalls/linux.Connect() pkg/sentry/syscalls/linux/sys_socket.go:286 +0x251 Previous read at 0x00c0006aa2b8 by goroutine 270: pkg/sentry/socket/unix/transport.(baseEndpoint).Connected() pkg/sentry/socket/unix/transport/unix.go:789 +0x42 pkg/sentry/socket/unix/transport.(connectionedEndpoint).State() pkg/sentry/socket/unix/transport/connectioned.go:479 +0x2f pkg/sentry/socket/unix.(socketOpsCommon).State() pkg/sentry/socket/unix/unix.go:714 +0xc3e pkg/sentry/socket/unix.(socketOpsCommon).SendMsg() pkg/sentry/socket/unix/unix.go:466 +0xc44 pkg/sentry/socket/unix.(SocketOperations).SendMsg() <autogenerated>:1 +0x173 pkg/sentry/syscalls/linux.sendTo() pkg/sentry/syscalls/linux/sys_socket.go:1121 +0x4c5 pkg/sentry/syscalls/linux.SendTo() pkg/sentry/syscalls/linux/sys_socket.go:1134 +0x87 Reported-by: syzbot+c2be37eedc672ed59a86@syzkaller.appspotmail.com PiperOrigin-RevId: 317236996
2020-06-18	iptables: remove metadata struct	Kevin Krakauer
	Metadata was useful for debugging and safety, but enough tests exist that we should see failures when (de)serialization is broken. It made stack initialization more cumbersome and it's also getting in the way of ip6tables. PiperOrigin-RevId: 317210653
2020-06-18	Acquire lock when accessing MultiDevice's cache in String().	Ting-Yu Wang
	PiperOrigin-RevId: 317180925
2020-06-18	Remove various uses of 'whitelist'	Michael Pratt
	Updates #2972 PiperOrigin-RevId: 317113059
2020-06-18	Support setsockopt SO_SNDBUF/SO_RCVBUF for raw/udp sockets.	Bhasker Hariharan
	Updates #173,#6 Fixes #2888 PiperOrigin-RevId: 317087652
2020-06-17	Implement Sync() to directories	Fabricio Voznika
	Updates #1035, #1199 PiperOrigin-RevId: 317028108
2020-06-17	Internal change.	gVisor bot
	PiperOrigin-RevId: 316973783
2020-06-17	Remove various uses of 'blacklist'	Michael Pratt
	Updates #2972 PiperOrigin-RevId: 316942245
2020-06-17	Refactor host.canMap.	Dean Deng
	Simplify the canMap check. We do not have plans to allow mmap for anything beyond regular files, so we can just inline canMap() as a simple file mode check. Updates #1672. PiperOrigin-RevId: 316929654
2020-06-17	Implement POSIX locks	Fabricio Voznika
	- Change FileDescriptionImpl Lock/UnlockPOSIX signature to take {start,length,whence}, so the correct offset can be calculated in the implementations. - Create PosixLocker interface to make it possible to share the same locking code from different implementations. Closes #1480 PiperOrigin-RevId: 316910286
2020-06-16	Correctly handle multiple resizings in pgalloc.findAvailableRange().	Jamie Liu
	PiperOrigin-RevId: 316778032
2020-06-16	Port aio to VFS2.	Nicolas Lacasse
	In order to make sure all aio goroutines have stopped during S/R, a new WaitGroup was added to TaskSet, analagous to runningGoroutines. This WaitGroup is incremented with each aio goroutine, and waited on during kernel.Pause. The old VFS1 aio code was changed to use this new WaitGroup, rather than fs.Async. The only uses of fs.Async are now inode and mount Release operations, which do not call fs.Async recursively. This fixes a lock-ordering violation that can cause deadlocks. Updates #1035. PiperOrigin-RevId: 316689380
2020-06-16	Miscellaneous VFS2 fixes.	Jamie Liu
	PiperOrigin-RevId: 316627764
2020-06-12	vfs2: implement fcntl(fd, F_SETFL, flags)	Andrei Vagin
	PiperOrigin-RevId: 316148074
2020-06-11	Add //pkg/sentry/fsimpl/overlay.	Jamie Liu
	Major differences from existing overlay filesystems: - Linux allows lower layers in an overlay to require revalidation, but not the upper layer. VFS1 allows the upper layer in an overlay to require revalidation, but not the lower layer. VFS2 does not allow any layers to require revalidation. (Now that vfs.MkdirOptions.ForSyntheticMountpoint exists, no uses of overlay in VFS1 are believed to require upper layer revalidation; in particular, the requirement that the upper layer support the creation of "trusted." extended attributes for whiteouts effectively required the upper filesystem to be tmpfs in most cases.) - Like VFS1, but unlike Linux, VFS2 overlay does not attempt to make mutations of the upper layer atomic using a working directory and features like RENAME_WHITEOUT. (This may change in the future, since not having a working directory makes error recovery for some operations, e.g. rmdir, particularly painful.) - Like Linux, but unlike VFS1, VFS2 represents whiteouts using character devices with rdev == 0; the equivalent of the whiteout attribute on directories is xattr trusted.overlay.opaque = "y"; and there is no equivalent to the whiteout attribute on non-directories since non-directories are never merged with lower layers. - Device and inode numbers work as follows: - In Linux, modulo the xino feature and a special case for when all layers are the same filesystem: - Directories use the overlay filesystem's device number and an ephemeral inode number assigned by the overlay. - Non-directories that have been copied up use the device and inode number assigned by the upper filesystem. - Non-directories that have not been copied up use a per-(overlay, layer)-pair device number and the inode number assigned by the lower filesystem. - In VFS1, device and inode numbers always come from the lower layer unless "whited out"; this has the adverse effect of requiring interaction with the lower filesystem even for non-directory files that exist on the upper layer. - In VFS2, device and inode numbers are assigned as in Linux, except that xino and the samefs special case are not supported. - Like Linux, but unlike VFS1, VFS2 does not attempt to maintain memory mapping coherence across copy-up. (This may have to change in the future, as users may be dependent on this property.) - Like Linux, but unlike VFS1, VFS2 uses the overlayfs mounter's credentials when interacting with the overlay's layers, rather than the caller's. - Like Linux, but unlike VFS1, VFS2 permits multiple lower layers in an overlay. - Like Linux, but unlike VFS1, VFS2's overlay filesystem is application-mountable. Updates #1199 PiperOrigin-RevId: 316019067
2020-06-11	Merge pull request #2863 from lubinszARM:pr_sndbuf	gVisor bot
	PiperOrigin-RevId: 315991648
2020-06-11	Don't copy structs with sync.Mutex during initialization	Fabricio Voznika
	During inititalization inode struct was copied around, but it isn't great pratice to copy it around since it contains ref count and sync.Mutex. Updates #1480 PiperOrigin-RevId: 315983788
2020-06-10	Remove duplicate colon from warning log.	Nicolas Lacasse
	doAction()->log.TracebackAll() will append a colon. PiperOrigin-RevId: 315842611
2020-06-10	Deleting the maxSendBufferSize from fs/host	Bin Lu
	When I do high-performance networking, the value of wmem_max is often set very high, specially for 10/25/50 Gigabit NIC. I think maybe this restriction is not suitable. Signed-off-by: Bin Lu <bin.lu@arm.com>
2020-06-10	Merge pull request #2711 from lubinszARM:pr_mmio	gVisor bot
	PiperOrigin-RevId: 315812219
2020-06-10	Merge pull request #2763 from ↵	gVisor bot
	gaurav1086:sentry_kernel_timekeeper_use_buffered_channel PiperOrigin-RevId: 315803553
2020-06-10	{S,G}etsockopt for TCP_KEEPCNT option.	Nayana Bidari
	TCP_KEEPCNT is used to set the maximum keepalive probes to be sent before dropping the connection. WANT_LGTM=jchacon PiperOrigin-RevId: 315758094
2020-06-10	socket/unix: handle sendto address argument for connected sockets	Andrei Vagin
	In case of SOCK_SEQPACKET, it has to be ignored. In case of SOCK_STREAM, EISCONN or EOPNOTSUPP has to be returned. PiperOrigin-RevId: 315755972
2020-06-10	Merge pull request #2787 from lubinszARM:pr_race_time	gVisor bot
	PiperOrigin-RevId: 315734425
2020-06-10	Redirect TODOs to more specific issues	Fabricio Voznika
	Closes #1623 PiperOrigin-RevId: 315681993
2020-06-09	sentry: use defer wg.Done() unconditionally	Gaurav Singh
	Signed-off-by: Gaurav Singh <gaurav1086@gmail.com>
2020-06-09	Implement flock(2) in VFS2	Fabricio Voznika
	LockFD is the generic implementation that can be embedded in FileDescriptionImpl implementations. Unique lock ID is maintained in vfs.FileDescription and is created on demand. Updates #1480 PiperOrigin-RevId: 315604825
2020-06-09	Merge pull request #2712 from lubinszARM:pr_sigfp_init	gVisor bot
	PiperOrigin-RevId: 315599736
2020-06-09	Merge pull request #2907 from lubinszARM:pr_minor	gVisor bot
	PiperOrigin-RevId: 315595602
2020-06-09	Don't WriteOut to readonly mounts	Fabricio Voznika
	When the file closes, it attempts to write dirty cached attributes to the file. This should not be done when the mount is readonly. PiperOrigin-RevId: 315585058
2020-06-09	Ensure pgalloc.MemoryFile.fileSize is always chunk-aligned.	Jamie Liu
	findAvailableLocked() may return a non-aligned FileRange.End after expansion since it may round FileRange.Start down to a hugepage boundary. PiperOrigin-RevId: 315520321
2020-06-09	minor change in kvm module for Arm64	Bin Lu
	Signed-off-by: Bin Lu <bin.lu@arm.com>
2020-06-09	initialize an empty fp state area for sentry on Arm64	Bin Lu
	We need to initialize an empty fp state area for the sentry. Signed-off-by: Bin Lu <bin.lu@arm.com>
2020-06-08	Combine executable lookup code	Fabricio Voznika
	Run vs. exec, VFS1 vs. VFS2 were executable lookup were slightly different from each other. Combine them all into the same logic. PiperOrigin-RevId: 315426443
2020-06-08	Implement VFS2 tmpfs mount options.	Jamie Liu
	As in VFS1, the mode, uid, and gid options are supported. Updates #1197 PiperOrigin-RevId: 315340510
2020-06-07	netstack: parse incoming packet headers up-front	Kevin Krakauer
	Netstack has traditionally parsed headers on-demand as a packet moves up the stack. This is conceptually simple and convenient, but incompatible with iptables, where headers can be inspected and mangled before even a routing decision is made. This changes header parsing to happen early in the incoming packet path, as soon as the NIC gets the packet from a link endpoint. Even if an invalid packet is found (e.g. a TCP header of insufficient length), the packet is passed up the stack for proper stats bookkeeping. PiperOrigin-RevId: 315179302
2020-06-05	Implement mount(2) and umount2(2) for VFS2.	Rahat Mahmood
	This is mostly syscall plumbing, VFS2 already implements the internals of mounts. In addition to the syscall defintions, the following mount-related mechanisms are updated: - Implement MS_NOATIME for VFS2, but only for tmpfs and goferfs. The other VFS2 filesystems don't implement node-level timestamps yet. - Implement the 'mode', 'uid' and 'gid' mount options for VFS2's tmpfs. - Plumb mount namespace ownership, which is necessary for checking appropriate capabilities during mount(2). Updates #1035 PiperOrigin-RevId: 315035352
2020-06-05	Add +checkescape annotations to kvm/ring0.	Adin Scannell
	This analysis also catches a potential bug, which is a split on mapPhysical. This would have led to potential guest-exit during Mapping (although this would have been handled by the now-unecessary retryInGuest loop). PiperOrigin-RevId: 315025106
2020-06-05	Use top-down allocation for pgalloc.	Adin Scannell
	This change has multiple small components. First, the chunk size is bumped to 1GB in order to avoid creating excessive VMAs in the Sentry, which can lead to VMA exhaustion (and hitting limits). Second, gap-tracking is added to the usage set in order to efficiently scan for available regions. Third, reclaim is moved to a simple segment set. This is done to allow the order of reclaim to align with the Allocate order (which becomes much more complex when trying to track a "max page" as opposed to "min page", so we just track explicit segments instead, which should make reclaim scanning faster anyways). Finally, the findAvailable function attempts to scan from the top-down, in order to maximize opportunities for VMA merging in applications (hopefully preventing the same VMA exhaustion that can affect the Sentry). PiperOrigin-RevId: 315009249
2020-06-05	Unshare files on exec	Andrei Vagin
	The current task can share its fdtable with a few other tasks, but after exec, this should be a completely separate process. PiperOrigin-RevId: 314999565
2020-06-05	Fix error code returned due to Port exhaustion.	Bhasker Hariharan
	For TCP sockets gVisor incorrectly returns EAGAIN when no ephemeral ports are available to bind during a connect. Linux returns EADDRNOTAVAIL. This change fixes gVisor to return the correct code and adds a test for the same. This change also fixes a minor bug for ping sockets where connect() would fail with EINVAL unless the socket was bound first. Also added tests for testing UDP Port exhaustion and Ping socket port exhaustion. PiperOrigin-RevId: 314988525
2020-06-05	Fix copylocks error about copying IPTables.	Ting-Yu Wang
	IPTables.connections contains a sync.RWMutex. Copying it will trigger copylocks analysis. Tested by manually enabling nogo tests. sync.RWMutex is added to IPTables for the additional race condition discovered. PiperOrigin-RevId: 314817019
2020-06-04	avoid runtime fails with missing stack maps in race mode on Arm64	Bin Lu
	In race mode, when calling the go function in asm code, there will be an missing stack maps issue. The root cause is: The function of 'muldiv64' has a non-empty frame, so it needs stack maps for locals, for which the macro NO_LOCAL_POINTERS will do. Also, the macro GO_ARGS can covers arguments. Signed-off-by: Bin Lu <bin.lu@arm.com>
2020-06-03	Pass PacketBuffer as pointer.	Ting-Yu Wang
	Historically we've been passing PacketBuffer by shallow copying through out the stack. Right now, this is only correct as the caller would not use PacketBuffer after passing into the next layer in netstack. With new buffer management effort in gVisor/netstack, PacketBuffer will own a Buffer (to be added). Internally, both PacketBuffer and Buffer may have pointers and shallow copying shouldn't be used. Updates #2404. PiperOrigin-RevId: 314610879