daemon #
Start a program in a session of its own, without forking the caller.
Daemon.spawn starts an executable as the leader of a new session and process
group with no controlling terminal, so it survives the caller's exit, the
caller's process group being signalled, and the caller's terminal going away.
The child comes from posix_spawn(3) with POSIX_SPAWN_SETSID, or from
vfork(2) and execve(2) where that call cannot keep the caller's other
descriptors out of the child, so the caller is never duplicated. A domain is
OCaml 5's unit of parallelism: a system thread with its own minor heap that
runs OCaml code in parallel with the others. fork(2) copies only the calling
thread, so a child forked while another domain runs could inherit the heap or
a lock in the middle of that domain's update; OCaml 5's Unix.fork therefore
fails for the rest of a process's life once any domain has been spawned, and
duplicating a multi-threaded runtime is a hazard where it is still allowed.
Requires OCaml >= 5.1. Supported on macOS and on Linux, the platforms which can keep the caller's other descriptors out of the child.
Security #
- The child receives the three descriptors it is given as its standard input,
output and error, plus explicitly declared
extra_fds, and no others, whatever the close-on-exec state of the caller's descriptors and whatever another domain opens while the call runs. Each mechanism settles the set to close in the child, whose descriptor table is the copy taken when it was made:POSIX_SPAWN_CLOEXEC_DEFAULTon macOS,posix_spawn_file_actions_addclosefrom_npon glibc 2.34 and later, andclose_range(2)between avfork(2)and theexecve(2)on any other Linux libc, whereposix_spawnhas no file action which closes a range and no hook for one. A platform with none of the three fails to build rather than spawn something less confined. - Signal dispositions are reset to default and the signal mask emptied, so an ignored or blocked signal in the caller does not silently disable the same signal in the new program.
- A string carrying a NUL byte is refused before anything reaches the C boundary, where it would be read as the end of a shorter argument than the caller wrote.
- The child is a session leader with no controlling terminal, so it cannot be reached by a terminal's job-control signals and cannot acquire the caller's terminal.
Install #
$ opam install daemon
If opam cannot find the package, add the overlay repository first:
$ opam repo add samoht https://tangled.org/gazagnaire.org/opam-overlay.git
$ opam update
$ opam install daemon
Usage #
(* Start a launcher which outlives this process, with its output in a log. *)
let start ~log exe =
let input = Unix.openfile "/dev/null" [ Unix.O_RDONLY ] 0 in
let output =
Unix.openfile (Fpath.to_string log)
[ Unix.O_WRONLY; Unix.O_CREAT; Unix.O_APPEND ]
0o600
in
let started =
Daemon.spawn ~stdin:input ~stdout:output ~stderr:output exe
[ Fpath.to_string exe; "--serve" ]
in
(* The child has its own copies; these are the caller's. *)
Unix.close input;
Unix.close output;
match started with
| Ok pid -> Fmt.pr "launcher running as pid %d@." pid
| Error (`Msg message) -> prerr_endline message
API #
Daemon.spawn ?cwd ?env ~stdin ~stdout ~stderr exe argvstartsexewith the whole argument vectorargv(its head isargv.(0)) and is its pid.cwdis the child's working directory, the caller's by default;envis the child's whole environment asKEY=VALUEbindings, the caller's by default. A program which cannot be executed (ENOENT,EACCES,ENOEXEC) is an error returned by the call, and no child is left to exit with 127. No exception escapes.Daemon.wait_exited pidwaits for the pidspawnreturned to end without reaping it, so a caller that has killed the process grouppidstill names (kill (-pid)) can do so before the number is freed by a reap, then reap the leader itself withUnix.waitpid.
See lib/daemon.mli for the full documentation.
From Eio #
daemon.eio is the same call over Eio flows. Daemon_eio.spawn unwraps the
descriptor behind each flow for the duration of the call, so a program which
opens its log through an Eio.Path capability never handles a descriptor
itself. A stream not given is /dev/null, opened for the call and closed once
the child holds its copy. A flow with no descriptor to give (a buffer, a mock,
a TLS flow) is refused before anything is spawned. The descriptor a
flow carries is put in blocking mode first, since the child's copy shares the
flag and the eio_posix backend opens every file and both ends of a pipe with
O_NONBLOCK (eio_linux, only its pipes). It stays cleared afterwards, spawn or
no spawn: giving the flag back would give it back to the child.
(* The same launcher, its console opened through an Eio capability. *)
let start_eio ~sw ~log exe =
let console =
Eio.Path.open_out ~sw ~append:true ~create:(`If_missing 0o600) log
in
match
Daemon_eio.spawn ~stdout:console ~stderr:console exe
[ Fpath.to_string exe; "--serve" ]
with
| Ok pid -> Fmt.pr "launcher running as pid %d@." pid
| Error e -> Fmt.epr "%a@." Daemon_eio.pp_error e
Daemon_eio.wait_exited pid is Daemon.wait_exited from a fiber. The pure
call is a blocking waitid(2) and holds the whole domain, every other fiber
and timer included, for as long as the child lives. This one runs it on a
system thread, so only the calling fiber waits. It leaves the child unreaped just the
same, and cannot be cancelled once the wait has begun.
See lib/eio/daemon_eio.mli for the full
documentation.
Related work #
Unix.create_process, Lwt_process and Eio.Process start a child in the
caller's session and process group. Taking a new session on top of one of them
means forking and calling setsid(2) in the child, which is the shape this
library replaces: posix_spawn(3) asks the kernel for both at once, so the
call is available to a process which has spawned a domain and needs no child
of its own to run OCaml code.
License #
ISC. See LICENSE.md.
Owned process groups #
Daemon.Managed.run starts a bounded command through the installed
daemon-supervisor companion. Daemon_eio.Managed.run adds fiber cancellation:
leaving the fiber closes its ownership pipe and joins cleanup. Abrupt owner
exit also closes that pipe, so it does not depend on an OCaml finalizer.
The supervisor keeps a live process-group anchor until it sends TERM followed by KILL, then reaps the anchor. It never signals a PID recovered from disk. An optional descriptor-owned lease stays in the supervisor until the group has stopped; commands do not inherit either the lease or the ownership pipe. A group the kernel cannot terminate retains its guardian and lease, while the caller receives a bounded refusal. The protocol uses a bounded JSON outcome.
Install the companion with the library, or set DAEMON_SUPERVISOR to its path
when running from a build tree. This manages cooperating process groups: a
command which deliberately creates a new session needs an external security
container. A supervisor that dies leaves the group to its anchor, which reads
the end of file of a pipe only the supervisor writes, stops the group with TERM
then KILL, and ends. A TERM sent to the anchor from outside stops the command
the same way, and the run reports how it ended.