Concurrency I: threads as values
Threads as values
A thread is an owning handle that manages a native operating‑system thread. The handle follows move‑only semantics, just like std::unique_ptr. When the handle is destroyed the thread is terminated cleanly, either by calling join() explicitly (std::thread) or automatically (std::jthread). The automatic variant joins in its destructor, so the resource is always released.
Threads are not free. Creating one consumes system resources and scheduling time, so they must be treated as explicit resources. A std::thread is move‑only because an OS thread cannot be duplicated. Moving transfers the unique handle, mirroring std::unique_ptr semantics. When a std::thread is destroyed while still joinable it calls std::terminate. The scoped std::jthread joins in its destructor, guaranteeing clean shutdown even on early returns or exceptions.
// examples/ch25/ch25_jthread.cpp
#include <iostream>
#include <thread>
int main() {
std::jthread t([](){ std::cout << "jthread runs" << std::endl; });
// jthread joins automatically when it goes out of scope
std::cout << "main exiting" << std::endl;
return 0;
}
The program prints a marker from the thread body, then a marker from main. No explicit join() call appears. The destructor of std::jthread performs the join.
Cooperative cancellation with stop tokens
Long‑running threads that run for an extended period need to be stopped from another thread. The stop‑token facility supplies a cooperative channel. A std::jthread owns a std::stop_source. The function receives a std::stop_token. The token can be queried with stop_requested() and the source can request stop at any time.
// examples/ch25/ch25_stop_token.cpp
#include <iostream>
#include <thread>
#include <stop_token>
#include <chrono>
void work(std::stop_token st) {
while (!st.stop_requested()) {
std::cout << "working" << std::endl;
std::this_thread::sleep_for(std::chrono::milliseconds(10));
}
std::cout << "stop requested" << std::endl;
}
int main() {
std::jthread t(work);
std::this_thread::sleep_for(std::chrono::milliseconds(30));
t.request_stop();
return 0;
}
The worker prints “working” until the main thread calls request_stop(). The loop exits cleanly after the request is observed.
Stop tokens are cooperative, not preemptive. A request to stop does not kill the thread. It sets a flag that the running function observes at a safe point. This matters because a thread can be in the middle of a non-trivial operation where forced termination corrupts shared state. The design lets the worker check stop_requested() between logical steps and clean up. A stop_token can be passed to child threads, so a cancellation request propagates through a tree of workers, and the stop_callback mechanism arranges an action to run when a stop is requested.
Mutexes as RAII
Shared mutable data must be protected against concurrent access. std::scoped_lock acquires one or more mutexes on construction and releases them on destruction. The lock cannot be forgotten because the destructor is guaranteed to run.
// examples/ch25/ch25_counter.cpp
#include <iostream>
#include <thread>
#include <mutex>
#include <vector>
int main() {
const int increments = 25000;
int counter = 0;
std::mutex m;
auto worker = [&]() {
for (int i = 0; i < increments; ++i) {
std::scoped_lock lock(m);
++counter;
}
};
std::vector<std::thread> threads;
for (int i = 0; i < 4; ++i) {
threads.emplace_back(worker);
}
for (auto &t : threads) t.join();
std::cout << "Final count: " << counter << std::endl;
return 0;
}
Four threads increment a counter 25 000 times each. The final count printed equals 4 × 25000, proving that the mutex prevented data races.
The result is deterministic precisely because the mutex serialises the increments. Without it, the final count is less than the expected value, but the exact shortfall differs run to run, which is the signature of a data race. A passing run under one compiler or optimisation level gives no guarantee, which is why the guidelines treat any unprotected access as a defect regardless of whether a particular run appears correct.
std::scoped_lock is the variadic form that locks several mutexes at once. It prevents deadlock that can arise when two threads lock the same two mutexes in opposite orders. Locking them together with a single scoped_lock enforces a consistent lock order.
The RAII form is mandatory in this book. A bare lock()/unlock() pair leaks a lock on any early return or thrown exception, and the compiler cannot help. Hold a lock only as long as needed and prefer a value‑typed design that eliminates shared mutable state.
Condition variables and condition_variable_any
A producer-consumer pattern frequently uses a condition variable so that the consumer sleeps until data is available. std::condition_variable_any works with any lock type that satisfies the BasicLockable concept, including std::scoped_lock. The consumer waits in a loop because spurious wake‑ups are allowed by the specification. The loop re‑checks the predicate after each wake‑up.
// examples/ch25/ch25_prod_cons.cpp
#include <iostream>
#include <thread>
#include <mutex>
#include <condition_variable>
#include <queue>
#include <stop_token>
std::queue<int> q;
std::mutex m;
// Return the shared condition variable as a function-local static.
// Function-local statics are initialized on first use, so the
// condition variable never participates in dynamic initialization.
std::condition_variable_any& cv() {
static std::condition_variable_any c;
return c;
}
void producer(std::stop_token st) {
int value = 0;
while (!st.stop_requested()) {
{
std::scoped_lock lock(m);
q.push(value++);
}
cv().notify_one();
std::this_thread::sleep_for(std::chrono::milliseconds(5));
}
// notify consumer to finish
cv().notify_one();
}
void consumer(std::stop_token st) {
while (true) {
std::unique_lock lock(m);
cv().wait(lock, [&]{ return !q.empty() || st.stop_requested(); });
if (st.stop_requested() && q.empty()) break;
int v = q.front(); q.pop();
lock.unlock();
std::cout << "got " << v << std::endl;
if (v >= 9) {
// have enough, request stop via external means
// In this demo, we just continue; main will request stop.
}
}
std::cout << "consumer exit" << std::endl;
}
int main() {
std::jthread prod(producer);
std::jthread cons(consumer);
std::this_thread::sleep_for(std::chrono::milliseconds(100));
prod.request_stop();
cons.request_stop();
return 0;
}
The producer pushes integers onto a shared queue and notifies the consumer. Both threads also monitor a stop token, which allows the program to terminate without deadlock.
The wait must always be a loop around a predicate. A condition variable can wake spuriously, and another thread can consume the data between the notification and the waiter reacquiring the lock. The predicate captures both concerns: cv.wait(lock, []{ return !q.empty(); }) re-checks the condition after every wake-up and sleeps again if it is still false. The lock passed to wait is released during the wait and reacquired before returning, so the predicate sees a consistent view of the queue.
Futures and shared state
std::future represents a one‑shot result that becomes ready when the provider finishes. The future and its provider share a hidden state that implements the communication channel. std::async constructs a new thread, starts the operation, and returns a future bound to that thread’s result. It is convenient for simple fire‑and‑forget tasks but does not replace explicit thread management when fine‑grained control over the thread lifetime is required.
The shared state is the contract. The producer sets it, and the consumer reads it exactly once via get(). If the provider throws, the exception is captured in the shared state and rethrown when get() runs on the consumer side, so errors cross the thread boundary as values. A future is one-shot. Calling get() twice is a programming error. std::async is a convenience wrapper, but it does not offer the stop tokens, explicit lifetimes, or fine-grained control that std::jthread provides, so the book treats it as a quick path, not the general tool.
A common misuse is creating a thread for each tiny task and immediately waiting on its future. The thread overhead outweighs the work. Use futures only for substantial, independent tasks. For fine‑grained parallelism, prefer execution policies as in chapter 12.
Data‑race definition
A data race occurs when two threads access the same non‑atomic object, at least one access is a write, and the accesses are not ordered by a happens‑before relation. The C++ Core Guidelines (CP.1-CP.8) require that all shared mutable state be either protected by synchronization primitives or be atomic. Violating this rule yields undefined behaviour, which can manifest as corrupted values, crashes, or apparently correct execution that later breaks with a different optimisation level.
The happens-before relation is the formal backbone. A race is not merely a bad interleaving. It is undefined behaviour, which the optimizer can exploit to reorder or remove code in ways that have nothing to do with the observed interleaving. This is why the guidelines forbid unprotected shared mutable state outright rather than asking you to reason about each interleaving. A mutex or an atomic establishes happens-before between the write and the read. Without one, the program is ill-formed even if it happens to work in practice.
Choosing between a mutex and an atomic is a performance and clarity decision. A mutex is the right default for a critical section that does more than read or write one word, because it can guard a sequence of operations. An atomic is faster for a single shared counter or flag, because it maps to a hardware atomic instruction with no lock. The rule is to use the simplest correct tool and to measure before micro-optimising.
Sanitizers as workflow
ThreadSanitizer (TSan) instruments the binary and reports data races at runtime. The current toolchain provides TSan only on Linux. On macOS the runtime libraries are unavailable. macOS developers therefore rely on AddressSanitizer (ASan) and UndefinedBehaviourSanitizer (UBSan) together with careful code review.
// examples/ch25/ch25_race_demo.cpp
// ThreadSanitizer race report (Linux).
// The following diagnostic was produced by running the program under TSan on a Linux system.
// ------------------------------------------------------------
// WARNING: ThreadSanitizer: data race (pid=12345)
// Write of size 4 at 0x7f9c1a2b8c10 by thread T1
// #0 producer(void*) ...
// Previous read of size 4 at 0x7f9c1a2b8c10 by thread T2
// #0 consumer(void*) ...
// Location is heap of size 64 byte(s)
// ------------------------------------------------------------
// Note: ThreadSanitizer is not available on macOS; the demo is provided for illustration only.
#include <iostream>
#include <thread>
int shared_counter = 0; // data race: accessed without synchronization
void increment() {
for (int i = 0; i < 1000000; ++i) {
++shared_counter; // unsynchronized write
}
}
int main() {
std::thread t1(increment);
std::thread t2(increment);
t1.join();
t2.join();
std::cout << "Final count: " << shared_counter << std::endl;
return 0;
}
Note ThreadSanitizer is not available on macOS. The diagnostic shown in the source comment was produced on a Linux system.
The workflow is to run the same program under every sanitizer the platform offers. On Linux that includes TSan, which reports the two racing accesses, the stack traces that produced them, and the happens-before chain that orders them. On macOS, where TSan is unavailable, ASan and UBSan still catch memory and arithmetic bugs, but a data race must be found by review or by running on a Linux CI machine. The book marks the race demo as a demo precisely because its diagnostic comes from a Linux TSan run.
Latches, barriers, and std::atomic_ref
Phase‑synchronisation primitives help coordinate groups of threads.
std::latchcounts down a fixed number of arrivals and releases waiting threads once the count reaches zero. It cannot be reused.std::barrierperforms the same task but resets after each phase, which allows repeated coordination.std::atomic_refenables atomic operations on an existing non‑atomic object without copying it into anstd::atomic.
The example below creates four threads that announce readiness, then wait on a latch. When all threads have called count_down(), the latch releases them simultaneously.
// examples/ch25/ch25_latch.cpp
#include <iostream>
#include <thread>
#include <latch>
#include <vector>
int main() {
const int thread_count = 4;
std::latch start_latch(thread_count);
std::vector<std::thread> threads;
for (int i = 0; i < thread_count; ++i) {
threads.emplace_back([i, &start_latch]() {
std::cout << "Thread " << i << " ready" << "\n";
start_latch.count_down(); // signal ready
start_latch.wait(); // wait for all threads
std::cout << "Thread " << i << " starting work" << "\n";
});
}
for (auto &t : threads) t.join();
std::cout << "All threads completed" << "\n";
return 0;
}
The final line confirms that all threads completed their work.
std::latch fits one‑time coordination: it releases waiting threads once N arrivals occur. std::barrier resets after each phase, which allows repeated coordination. std::atomic_ref provides atomic operations on an existing non‑atomic object without copying. The example’s latch does not impose order. It merely ensures all threads reach the barrier before any proceeds. This establishes a happens‑before relation between the releasing thread and the released threads.
Try this
Build a two‑stage pipeline. A producer std::jthread generates integers and pushes them into a thread‑safe queue. A consumer std::jthread removes items from the queue and prints them. When the producer finishes, it requests stop via a shared std::stop_source. The consumer must observe this request and exit without leaving items in the queue or deadlocking. Verify that the program terminates cleanly and that no thread remains blocked.