Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Preface

This book teaches C++ as if it were always C++26, using only a small allow‑list of legacy constructs taught explicitly: raw pointers and new/delete for RAII, virtual functions for type erasure, C strings and errno for C interop, and SFINAE for reading existing code.

The book has four parts and a set of appendices. Part I covers the language. Part II covers the standard library. Part III covers generic and compile‑time programming. Part IV covers systems topics. The appendices hold reference material, including a feature table and a Core Guidelines index.

Who this book is for

You are an experienced programmer who already knows C, Rust, Lisp, and Prolog, and you have never written C++. The book assumes you understand functions, structs, loops, ownership, and metaprogramming at a high level. It never wastes a sentence describing a basic function. Instead, every paragraph highlights what is specific to C++, or how C++ differs from the languages you already master.

Each chapter assumes you can read code in at least one of these languages. It does not re‑teach loops or structs. It focuses on the C++ way to express an idea you already know. You will learn the vocabulary of modern C++, the ownership model, and the tools that verify the examples.

How the book reads

The book is a tour, not a reference. Each chapter introduces a cluster of related ideas, connects them to what came before, and ends with a single challenge under a “Try this” heading. There is no solutions appendix. The examples are real files, compiled and run by the toolchain that ships with the book, so what you read is exactly what the compiler accepted.

The tour is cumulative. Later chapters assume the vocabulary of earlier ones. The capstone in Part III ties the ideas together with a constexpr SQL interpreter. You can read the chapters in order, or jump to a topic and follow its cross‑references back.

The toolchain

The book is built with a Nix dev shell that pins LLVM 22.1.8, its libc++, and the Clang tools. nix develop gives you the whole environment. scripts/verify.sh runs the complete acceptance pipeline. The pipeline builds the book and verifies that every example is correct. It follows this order:

  1. scripts/verify.sh enters the Nix dev shell and runs scripts/verify-inner.sh. The inner script checks that clang++ is version 22 or newer, because the lifetime‑safety analysis requires it.
  2. CMake configures the build with cmake --preset dev. It probes the compiler for the lifetime‑safety flag and assembles the warning set.
  3. cmake --build build compiles every example with the warning set, -Werror, and clang‑tidy. A warning becomes a hard error, so a clean example stays clean.
  4. CTest runs each compiled example under the sanitizers. Each book_example test must match its EXPECT regular expression.
  5. mdbook build renders the book. It includes each example file directly from examples/, so the printed code is the compiled code.

A feature that is new in C++26 but not yet implemented by this compiler is taught prose‑first and marked with a “Not yet deployable” callout, so the book never presents code it cannot verify.

A note on style

The code follows the C++ Core Guidelines. Types and functions use snake_case, functions declare their return type before their name, and errors are reported through exceptions or std::expected, never through errno or out‑parameters. Wherever a rule comes from the Core Guidelines, it is cited inline, for example (CG F.15), and collected in Appendix C. The prose emphasizes clarity, brevity, and direct comparison to your existing language experience.

The compiler is a committee

The thesis

This book treats the C++ toolchain as a single committee: clang++ front‑end, clang‑tidy, path‑sensitive static analyzer, address and undefined‑behavior sanitizers, and clangd. The clang++ front‑end, the clang‑tidy lint suite, the path‑sensitive static analyzer, the address and undefined‑behavior sanitizers, and the clangd language server together form the compiler committee. A diagnostic that originates from any of these components is a violation of the language law, not a mere suggestion. The Core Guidelines describe rules that the committee intends to make machine enforceable, and the committee enforces them.

clang++        front-end compiler
clang-tidy     lint suite
clang-analyzer path-sensitive static analyzer
ASan / UBSan   runtime sanitizers
clangd         language server

The mental model mirrors the Rust experience where the borrow checker, clippy linter, and cargo‑test runner cooperate to keep code correct. In the Rust ecosystem the borrow checker enforces lifetimes, clippy provides style and correctness lints, and cargo test runs the unit test suite.

In the C++ world those responsibilities are split across clang++ flags, clang‑tidy checks, the static analyzer and sanitizers, but the book treats them as one cohesive compiler surface.

Getting the toolchain

The canonical way to obtain the exact toolchain used by the book is to invoke the Nix flake that lives at the repository root. Running nix develop spawns a development shell that contains the pinned versions:

  • LLVM 22.1.8 provides clang++, clang‑tidy, clang‑analyzer and clangd.
  • CMake 4.3.4 drives the builds.
  • mdBook 0.5.4 renders the final HTML.

When Nix is unavailable the book falls back to system packages. On macOS the user can install the same LLVM version with brew install llvm. On Linux the user can follow the instructions at apt.llvm.org to obtain a recent clang. The script scripts/verify.sh builds every example, runs the tests, and builds the book. The script runs CMake to configure the build, builds all book_example targets, executes each test with CTest, and finally calls mdbook build to ensure the markdown renders without missing includes. The development shell also sets CC=clang and CXX=clang++. CMake locates clang‑tidy on its own with find_program. The same environment provides clangd for IDE integration.

Hello, world, honestly

The book writes #include <print> because the header <print> is the modern way to reach std::println. The directive #include is the build model the book uses, because the classic preprocessor workflow is what real multi‑file projects run today. The preprocessor expands #include before compilation.

The preprocessor gets full treatment in chapter 23, and modules in chapter 24.

std::println replaced std::cout and printf as the default output mechanism in C++23. It checks format strings at compile time, writes directly to stdout, and returns void. The older facilities still work, but the book uses the modern one from the first example so the reader never has to unlearn a habit.

#include <print>

int main() {
    std::println("hello, world");
}

Compiling the file by hand looks like this:

clang++ -std=c++26 -Wall -Wextra -Wpedantic -c ch01_hello.cpp
clang++ -std=c++26 -Wall -Wextra -Wpedantic ch01_hello.o -o ch01_hello

The short command line shows the language version flag, the three warning groups that the book treats as law, and the absence of any additional options. The resulting binary prints hello, world on standard output. The header <print> became part of the standard library in C++23 and is fully supported by the libc++ shipped with the pinned toolchain.

Warnings are laws

The next example illustrates a classic lifetime violation. The file returns a pointer to a local variable, a pattern that the compiler can diagnose.

// DEMO: deliberately wrong. This file exists to produce a diagnostic.
// The function returns a pointer to a dead local. The lifetime-safety
// analysis and the plain -Wreturn-stack-address warning both catch it.
//
//   examples/ch01/ch01_dangling.cpp:9:13: warning: address of stack memory
//   associated with local variable 'x' returned [-Wreturn-stack-address]
//       9 |     return &x;
//         |             ^

int* bad() {
    int x = 42;
    return &x;
}

int main() {
    int* p = bad();
    return *p;
}

The diagnostic that appears in the comment of the example appears again here verbatim:

examples/ch01/ch01_dangling.cpp:9:13: warning: address of stack memory associated with local variable 'x' returned [-Wreturn-stack-address]
    9 |     return &x;
      |             ^

The -Wreturn-stack-address check emits the warning, and the check is part of the standard warning set. The book treats this warning as a language law. The book compiles its real examples with -Werror, which turns every warning into a hard error, and encourages the reader to do the same. The deliberately wrong demos keep their warning visible to serve as the exhibit, as described under book_example and book_demo in the section ## CMake As The Contract. The larger ambition of the committee is to provide a unified lifetime‑safety profile. The flag that represents this profile in current clang builds is -Wlifetime-safety. On the pinned clang 22.1.8 this flag does not exist directly. Instead the compiler accepts the experimental spelling -Xclang -fexperimental-lifetime-safety, but it produces almost no diagnostics on the cases shown in chapter 7. The book therefore relies on the diagnostics that do fire in the shipped toolchain, while still naming the lifetime‑safety analysis as the law that the committee intends to enforce.

Not yet deployable. The lifetime‑safety analysis is young. The flag -Wlifetime-safety will become a full diagnostic source in future clang releases. Until then the analysis is visible only through the experimental option and the traditional warnings such as -Wreturn-stack-address. The flag -Wuninitialized also catches use of a variable before a value stores into it, and clang enables it by default.

Sanitizers as Compile Options

Dynamic checking is another member of the committee. The address sanitizer (ASan) and the undefined‑behavior sanitizer (UBSan) detect errors at runtime. The following example purposely reads past the end of a std::vector to trigger an ASan report.

// DEMO: deliberately wrong. This file exists to produce a sanitizer report.
// Run under AddressSanitizer, the read past the end of the vector storage
// reports heap-buffer-overflow:
//
//   ==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x6020000000fc
//   READ of size 4 at 0x6020000000fc thread T0
//       #0 ... std::__1::__format::__create_format_arg ... format_arg_store.h:191
//       ...
//       #8 ... in main ch01_overflow.cpp:13

#include <print>
#include <vector>

int main() {
    std::vector<int> v{1, 2, 3};
    volatile std::size_t past_the_end = v.size(); // volatile: defeat constant folding
    std::println("{}", v.data()[past_the_end]);
}

The report pasted in the file appears again exactly:

==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x6020000000fc
READ of size 4 at 0x6020000000fc thread T0
    #0 ... std::__1::__format::__create_format_arg ... format_arg_store.h:191
    ...
    #8 ... in main ch01_overflow.cpp:13

The build compiles the program with the sanitizer flags -fsanitize=address,undefined. The static members of the committee (clang‑tidy and the static analyzer) catch many issues at compile time, while sanitizers catch the remaining runtime bugs. The undefined‑behavior sanitizer reports a range of UB such as signed integer overflow, shift overflow and null‑pointer dereference. It inserts runtime checks that abort execution as soon as the offending operation occurs. The instrumentation stays in the finished binary: the address sanitizer roughly doubles runtime and memory use in typical programs, so sanitizer builds are a test configuration, not the shipping binary. The book treats both classes of diagnostics as equally binding.

The Static Analyzer

The clang static analyzer runs a path‑sensitive analysis that discovers bugs that are plain warnings miss. The command scan-build clang++ … invokes the analyzer and prints a concise HTML report. Integrated development environments that use clangd can invoke the analyzer on the fly. An example that the analyzer catches but the usual warnings miss is a use‑after‑free of a heap object passed through a function pointer. The analyzer reports the error with a trace that shows the allocation, the free, and the later dereference. The analyzer works by symbolically executing the control‑flow graph of each function and exploring the paths that reach each statement. It can track heap allocations across function boundaries and can flag memory leaks even when the program frees memory in a different function. The analysis is conservative. It reports false positives on some code patterns, and a reported error is still worth the reader’s time.

The Rulebook

The committee’s rulebook lives in the repository root as .clang‑tidy. The file appears verbatim below.

# The compiler committee's rulebook. Every book_example target builds with
# these checks; ch01 quotes this file verbatim. clang-tidy runs as part of
# compilation (see cmake/BookExample.cmake), so every retained check must stay
# silent on the examples.
#
# Stylistic checks are pruned: they either argue with the book's teaching
# style or fire on idiomatic example code. The retained correctness checks
# (bugprone, cppcoreguidelines, clang-analyzer, and the rest) stay on so a
# real defect in an example is still caught.
#
# Disabled checks are deliberate:
# - *avoid-magic-numbers: teaching examples are full of small literal values
#   whose names would add noise, not safety.
# - modernize-use-trailing-return-type: this book uses leading return types.
# - bugprone-exception-escape: every example main() may throw from std::println;
#   this check would flag all of them for no safety gain.
# - The stylistic checks below are pruned for the same reason: they fight the
#   book's style rather than find bugs.
Checks: >
  bugprone-*,
  cppcoreguidelines-*,
  modernize-*,
  performance-*,
  readability-*,
  clang-analyzer-*,
  -readability-identifier-length,
  -readability-braces-around-statements,
  -readability-named-parameter,
  -readability-implicit-bool-conversion,
  -readability-math-missing-parentheses,
  -readability-simplify-subscript-expr,
  -readability-isolate-declaration,
  -readability-convert-member-functions-to-static,
  -readability-make-member-function-const,
  -readability-else-after-return,
  -readability-container-data-pointer,
  -modernize-use-designated-initializers,
  -modernize-use-nodiscard,
  -modernize-use-std-numbers,
  -modernize-use-ranges,
  -modernize-avoid-bind,
  -modernize-type-traits,
  -modernize-use-constraints,
  -modernize-avoid-c-arrays,
  -performance-avoid-endl,
  -performance-unnecessary-value-param,
  -performance-inefficient-vector-operation,
  -bugprone-easily-swappable-parameters,
  -bugprone-crtp-constructor-accessibility,
  -bugprone-unsafe-functions,
  -cppcoreguidelines-avoid-non-const-global-variables,
  -cppcoreguidelines-pro-bounds-array-to-pointer-decay,
  -cppcoreguidelines-pro-bounds-pointer-arithmetic,
  -cppcoreguidelines-pro-bounds-constant-array-index,
  -cppcoreguidelines-pro-bounds-avoid-unchecked-container-access,
  -cppcoreguidelines-avoid-c-arrays,
  -cppcoreguidelines-owning-memory,
  -cppcoreguidelines-rvalue-reference-param-not-moved,
  -clang-analyzer-unix.Stream,
  -cppcoreguidelines-avoid-magic-numbers,
  -readability-magic-numbers,
  -modernize-use-trailing-return-type,
  -bugprone-exception-escape
WarningsAsErrors: '*'
HeaderFilterRegex: '.*'
FormatStyle: file

The configuration enables a broad set of checks grouped under the headings bugprone, cppcoreguidelines, modernize, performance, readability, and clang-analyzer. The bugprone group catches patterns that frequently cause bugs, for example misuse of std::move. The cppcoreguidelines group enforces the Core Guidelines, such as requiring initialization of all members. The modernize group suggests modern C++ idioms, for example replacing raw arrays with std::array. The performance group flags inefficient constructions such as unnecessary copies. The readability group encourages clear code, for example by preferring range‑based loops over index arithmetic. The clang-analyzer group runs the static analyzer during compilation. Three checks are deliberately disabled:

  • avoid-magic-numbers: teaching examples contain many small literal values and naming each one adds noise.
  • modernize-use-trailing-return-type: the book prefers leading return types for readability.
  • bugprone-exception-escape: every example main can throw from std::println. This check flags all of them without improving safety.

Only the enabled checks form part of the language law for the book. The HeaderFilterRegex: '.*' setting applies the checks to every file in the repository. The WarningsAsErrors: '*' setting promotes every retained clang-tidy warning to a hard error, so a book_example that trips a check fails the build. Compiler warnings become hard errors the same way through -Werror (see “Warnings are laws”). The deliberately-wrong book_demo targets run clang-tidy off, so their teaching diagnostics still reach the reader.

CMake As The Contract

The build model for each example is deliberately simple. A CMake macro called book_example wraps every compiled example. The macro adds the source file as a target, attaches the clang‑tidy checks, enables the sanitizers, and registers a test with CTest. The macro also adds the flag -Wlifetime-safety when the compiler supports it, so the lifetime profile runs for every example that uses a reference or pointer. When the test runs the program, CTest verifies the expected output or the presence of a diagnostic in the standard output. The macro hides the details of the build system. The reader only needs to understand that the committee compiles each example once, checks it, and exercises it with a test. No deeper CMake teaching is necessary at this point. A typical invocation looks like this in CMakeLists.txt:

book_example(ch01_hello.cpp EXPECT "hello, world")

The EXPECT argument tells CTest which regular expression to match on the program’s standard output. For the deliberately wrong demos the picture is different. They register a compile target but no test, because their exhibit is the compiler diagnostic itself, which the author pastes into the file’s header comment. The macro also defines a book_gap variant for features that are not yet supported by the pinned compiler. Those targets build only when the build passes the option BOOK_ENABLE_GAPS=ON.

Try This

Trigger a UBSan diagnostic

Write a program that adds two signed 32‑bit integers where the result overflows the representable range. Compile the program with the flags -fsanitize=address,undefined. Run the binary and observe that the undefined‑behavior sanitizer emits a diagnostic while the address sanitizer remains silent. Explain which member of the committee produced each message.

int main() {
    int a = 2'000'000'000;
    int b = 2'000'000'000;
    int sum = a + b;  // signed overflow, undefined behavior
    return sum == 0;
}

Make clang‑tidy flag a C‑style array

Create a tiny source file that defines a C‑style array of ten int values and fills it with a loop. Run clang‑tidy on the file using the repository configuration. The modernize-avoid-c-arrays check must produce a diagnostic suggesting the use of std::array. The program still compiles and runs, but the warning demonstrates how the committee enforces modern practices.

Values and functions

A C++ program is a graph of functions that exchange values. This book treats the function as the atomic unit of design: a named operation on values, written once and called from anywhere, not a method bolted to a class. Classes appear later and stay rare. If you come from Rust, read “function” as fn. If you come from C, returning rich values (tuples, structs, containers) is cheap and normal, so C’s out‑parameters and pointer‑returning habits fall away.

A function is one logical operation

A declaration tells the compiler a name’s type, and a definition supplies the body and must appear exactly once in the whole program (CG F.2). In a single-file example you write both at once. In a multi-file program the declaration lives in a header and the definition in one translation unit. Chapter 23 covers that split, and the one-definition rule that enforces it.

The guideline is deliberately narrow: one function does one logical operation.

Design each function to perform a single logical operation. This makes the function easy to test in isolation and reduces hidden coupling. For example, draw_triangle is a function, while draw_triangle_and_clear_screen_and_log is not. The compiler cannot enforce this rule. It relies on the programmer. When a function grows, extract helper functions so each piece remains testable and readable. Small functions can be inlined, eliminating call overhead.

Return by value is the default

C++ moves values, so returning a std::vector or a struct is not a copy of megabytes. The function builds the result in its own frame and the caller’s variable takes ownership of it through move, which for a temporary is elided to no copy at all. Write auto for the return type and let the compiler deduce it from the return expression:

#include <print>
#include <tuple>
#include <cmath>

// Solve a*x^2 + b*x + c = 0 and return the two roots.
// For the equation x^2 - 3*x + 2 = 0 the roots are 2 and 1.
std::tuple<double, double> solve_quadratic(double a, double b, double c) {
    double discriminant = b * b - 4.0 * a * c;
    double sqrt_disc = std::sqrt(discriminant);
    double root1 = (-b + sqrt_disc) / (2.0 * a);
    double root2 = (-b - sqrt_disc) / (2.0 * a);
    return {root1, root2};
}

int main() {
    // Structured binding decomposes the tuple into two distinct names.
    auto [root1, root2] = solve_quadratic(1.0, -3.0, 2.0);
    // Print exactly "roots: 2, 1".
    std::println("roots: {}, {}", root1, root2);
    return 0;
}

Multiple results come back as a std::tuple or, better, as a small named struct (CG F.21). You read them with a structured binding, which decomposes any aggregate (a tuple, a std::pair, an array, or a user-defined struct) into named locals in one line. The binding auto [root1, root2] is not a destructuring of a special tuple type, it is generic aggregate decomposition and you will lean on it for your own types in chapter 3.

Out-parameters (void solve(double*, double*)) are forbidden by the book’s style. They obscure the result in the argument list and break move semantics (CG F.20). Return the value.

Trailing return types, and why the book avoids them

The book writes the return type first, before the parameter list:

std::vector<int> make_squares(int n);

C++ also offers a trailing return type, written after the parameter list and preceded by ->:

auto make_squares(int n) -> std::vector<int>;

The book uses the leading form as its default. The trailing form reads right-to-left and hides the result behind the parameters, so it is not the book’s style. Yet the trailing form is required, or genuinely helpful, in a few real cases.

A trailing return type can name a parameter. In a leading return type the parameter names are not yet in scope, so you cannot write their types with decltype. The trailing form places the parameters first, which puts them in scope for the return type. This is the classic case where trailing is required:

template <typename T>
auto describe(T const& t) -> decltype(t.size());

The same scoping helps member function definitions. A trailing return type appears after the class name, so names looked up inside the class body are in scope. A leading return type sits before the class name and cannot see them:

struct Counter {
    std::size_t count() const;
};
auto Counter::count() const -> std::size_t; // return type sees the class

Returning a function pointer or an array reads badly with a leading type. The trailing form keeps the pointer or array syntax attached to the function name:

auto choose(int) -> int(*)(int); // returns a pointer to a function

Finally, the trailing form orders information the way the reader meets it: parameters first, result last. That ordering is the source of the “East End Functions” style. The book still prefers the leading return type for everyday code, and reserves the trailing form for the cases above where it is required or clearer.

#include <print>
#include <string_view>

// The return type is a trailing decltype of the parameter. Only the
// parameter name is in scope at that point, so this form is required.
template <typename T>
auto describe(T const& t) -> decltype(t.size()) {
    return t.size();
}

int main() {
    std::string_view greeting = "hello";
    std::println("size: {}", describe(greeting));
    return 0;
}

Overloading: one name, many shapes

Overloading lets several functions share a name, distinguished by argument types. C has no overloading. This is the first construct where C++ reads as a different language. The compiler picks the overload at compile time by best match: an exact match beats a promotion, a promotion beats a standard conversion, and a user-defined conversion is the last resort. The “best match wins” rule is all you need for everyday code. The full precedence machinery, including how templates enter the contest, waits for chapters 16 and 20.

#include <print>
#include <string_view>

void describe(int) {
    std::println("int overload");
}

void describe(double) {
    std::println("double overload");
}

void describe(std::string_view) {
    std::println("string overload");
}

int main() {
    describe(42);
    describe(3.14);
    describe(std::string_view{"hello"});
    return 0;
}

When the call sites are describe(42), describe(3.14), and describe("hello"), the compiler resolves each to a distinct function, and the test confirms the string overload fired. Overload resolution is zero-cost: the choice is made entirely before the program runs.

Overloading is the first place C++ diverges from C. C solves the same problem with name mangling: describe_int and describe_string are different names. C++ lets you use one name and trusts the compiler to pick. This is the same rule that templates extend in chapter 16, where the “shapes” become patterns and the compiler writes a new function for each type.

Default arguments state the common case

A parameter can carry a default used when the caller omits it. Defaults sit on the declaration and must be trailing: once a parameter has a default, every parameter to its right must too. A default encodes the usual call. The unusual call supplies the argument. Keep them for genuine common cases, not to merge several unrelated functions into one.

void greet(std::string_view name = "world") {
    std::println("hello, {}", name);
}

greet() prints hello, world. greet("Alice") prints hello, Alice. The default is compiled into each call site, so the two forms cost the same. (The multi-declaration rules for defaults across files belong to chapter 23.)

const is the default

const is a promise the compiler enforces. A const name cannot be rebound to a different value after initialization. Declare a local const unless it must change (CG ES.25). This is the same split as Rust’s let versus let mut, and the same discipline: start immutable, relax only where the algorithm demands it.

A const local variable binds once and never changes:

const double pi = 3.14159;
const int max_attempts = 3;

A const reference reads an object without copying or mutating it. Functions that merely read a value take it by const reference so they neither copy nor mutate (CG Con.1):

void show(const std::string& name) {
    std::println("{}", name);
}

The placement of const on a pointer decides what is fixed. A pointer to const (const T*) lets the pointer move but not the pointee. A const pointer (T* const) cannot be reseated, but the pointee stays mutable. A const pointer to const (const T* const) fixes both:

const int* p1;                // pointer to const, pointee read‑only
int* const p2 = &value;       // const pointer, cannot be reseated
const int* const p3 = &value; // const pointer to const

A const member function promises not to modify the object it runs on. The compiler checks this, so a const object can call it:

struct Counter {
    int value() const { return count_; }
    int count_ = 0;
};

A const return value protects the result from mutation. Returning const T by value is rare, because it blocks move semantics. The book returns plain values and reserves const for references and pointers:

const std::string label() const; // const return value, const member

A const function parameter promises the callee will not modify the argument. This is the default for read‑only parameters (chapter 8):

void draw(const Shape& shape);

Immutability‑by‑default is the cheapest correctness tool the language offers, and the book applies it everywhere.

Function naming and error handling

Name functions with a lower‑case verb phrase that describes the action, e.g., read_file or calculate_checksum. Avoid generic names. For recoverable errors return std::expected<T, E>. For unrecoverable failures throw an exception that includes a message identifying the failed operation.

References are aliases, not pointers

int& r = x; gives r a second name for x. Reads and writes through r reach the same object. A reference is not a pointer you must dereference, and it cannot be reseated after initialization. In Rust terms it is a reborrow.

A reference parameter hands the caller’s object to the callee, which can read or modify it. The full decision of what to pass (value, const reference, or std::span) is chapter 8. For now, the rule is that a reference means “I am” using your object,“ never “I own a copy.

Practical tips for writing functions

When designing a function, choose a clear verb‑phrase name, keep parameters short, pass large objects by const reference and small trivially copyable types by value, and return a value when needed.

Use void for functions that produce no observable result.

Common pitfalls

  • Do not return a reference to a local variable. The reference dangles after the function returns (see the WARNING above).
  • Do not store a pointer to a parameter that the caller can destroy before the function returns.
  • Avoid default arguments that depend on mutable global state, as they create hidden dependencies.
  • Keep overload sets small and well‑documented. Overloads that differ only by const qualification can be confusing.

Performance considerations

When a function returns a large object by value the compiler can apply return value optimization. This eliminates the temporary copy by constructing the result directly in the caller’s storage. Modern compilers can also apply named return value optimization when the return variable is a named local. The generated code frequently reduces memory traffic and improves cache usage.

If a function takes a large argument by value the copy can be expensive. Prefer a const reference for read‑only parameters. For parameters that the function will modify and then return, take the argument by value and move it back. This pattern enables the caller to pass a temporary without an extra copy. The move operation transfers ownership of the internal resources with minimal overhead.

Inlining small functions removes the call overhead entirely. The compiler decides whether inlining is beneficial based on the function size and the context of the call site. Functions that consist of a single return statement or a few arithmetic operations are prime candidates for inlining. Developers can hint at inlining with the inline keyword, but the final decision rests with the optimizer.

Do not include unnecessary branches inside hot loops. Branch prediction failures can stall the pipeline. When possible restructure code to keep the hot path straight. Use constexpr when the result can be computed at compile time. This moves work from runtime to compile time and can produce faster executables.

Profile the code with a sampling profiler to locate bottlenecks. Measure the impact of changes rather than assuming improvement. The guidelines in this chapter aim to produce clear, maintainable code, and the performance impact of each choice must be evaluated in the context of the whole program.

WARNING Never return a reference to a local variable. The local is destroyed when the function returns, and the caller receives a dangling name. Chapter 1 showed the compiler catching exactly this. Return-by-reference is reserved for assignment operators, which return *this (CG F.47).

Try this

Rewrite solve_quadratic to return a small user-defined struct { double first; double second; } instead of a std::tuple, and keep the call site a structured binding. In one sentence, state what this proves about structured bindings.

User-defined types: structs, sums, and expectations

Structs are product types

In C++ a struct groups a fixed set of named fields. The compiler automatically provides a default constructor, a copy constructor, a move constructor, and a trivial destructor when the fields themselves support those operations. This mirrors the mathematical notion of a product: a value of the struct type contains one value of each field.

struct point {
    double x;
    double y;
};

The declaration above yields a type that can be constructed with brace initialisation: point{1.0, 2.0}. No user-written constructor is required, which satisfies the Core Guidelines recommendation to prefer aggregates (C.20). The fields are public by default. The type behaves like a plain data carrier, unlike the historic C struct, which cannot contain member functions or const qualifiers.

Member functions can be added without sacrificing the aggregate property. A read-only member function is marked const so the compiler forbids it from modifying the object (C.2). The point example below adds double distance(const point&) const, which computes the Euclidean distance without changing *this and can therefore be called on a const point.

An aggregate also supports designated initialisers. The syntax point{.x = 1.0, .y = 2.0} names each field explicitly. The compiler checks the order and the names, so a typo in a field name fails to compile. Designated initialisers make the intent of each value clear at the call site. They also survive a change in field order without silently swapping values. This form is the clearest way to build a small data carrier.

The point type with its distance function is a complete example. The test verifies that the distance between two points is correct.

#include <print>
#include <cmath>

struct point {
    double x;
    double y;
    double distance(const point& other) const {
        double dx = x - other.x;
        double dy = y - other.y;
        return std::hypot(dx, dy);
    }
};

int main() {
    point p1{0.0, 0.0};
    point p2{3.0, 4.0};
    double d = p1.distance(p2);
    std::println("distance: {:.2f}", d);
    return 0;
}

Invariants and the small private we allow

A struct needs to enforce a relationship between its fields: an invariant. The book avoids class-based object orientation, yet a narrow use of private together with a public accessor is acceptable when the invariant cannot be expressed by the type system alone. For example, a normalized 3-D vector must always have length 1:

struct vec3 {
private:
    double x, y, z;
    vec3(double a, double b, double c) : x(a), y(b), z(c) {}
public:
    static vec3 make(double a, double b, double c) {
        double len = std::hypot(a, b, c);
        return vec3{a / len, b / len, c / len};
    }
    double length() const { return std::hypot(x, y, z); }
};

The private section prevents accidental mutation that can break the invariant, while the static factory make guarantees a correctly normalised instance. This follows Core Guidelines C.21 to keep data encapsulation minimal and prefer plain functions over heavy OO machinery. The invariant, that a vec3 constructed via make has length 1, is enforced by the factory. Direct use of the private constructor violates the contract, and the factory also checks for a zero‑length input and rejects it, avoiding a NaN result.

Because the type system cannot express this invariant, the responsibility rests with the factory and disciplined callers. This trades a small runtime check for the complexity of a large class hierarchy.

Defaulted comparison and the spaceship

C++20 introduced the three-way comparison operator <=>, commonly called the spaceship. When a struct’s fields already support equality and ordering, the compiler can generate all six relational operators automatically:

struct point {
    double x;
    double y;
    auto operator<=>(const point&) const = default;
};

Before C++20 each relational operator had to be written manually, a source of boilerplate and possible inconsistencies. The defaulted spaceship satisfies the Core Guidelines rule that the compiler must generate “obvious” functions (C.10, C.87). The generated operators perform a lexicographic comparison of the fields in declaration order, mirroring the behaviour of std::tie.

The spaceship is the last piece the point type needs. With it, the type supports equality, ordering, and structured bindings, all without a hand-written operator. The compiler generates the comparisons from the fields, and the fields are the data. This is the Rule of Zero applied to comparisons: the type declares its intent (= default), and the compiler does the work.

Structured bindings let a caller unpack a point into its fields in one statement. The form auto [px, py] = p binds px to p.x and py to p.y. The binding respects the declaration order of the fields, which matches the order that the spaceship compares. So the ordering of the fields has two consequences: it fixes the order of comparison and the order of the bindings. A reader who knows the field order can predict both. This consistency is why the book keeps field order stable across the type.

The spaceship also reports the strength of the comparison. The defaulted operator returns a comparison category, and the compiler derives it from the field types. For a point of two double fields the category is std::partial_ordering, because floating-point values do not always compare as equal to themselves. The caller rarely names this category directly, yet it governs how the generated operators behave. This detail matters when a struct mixes fields of different comparison strength.

std::variant as a sum type

A sum type represents a value that is exactly one of several alternatives. In C++ the standard library provides std::variant for this purpose. It stores a discriminated union together with a runtime tag, guaranteeing safe access. The expression-tree example below demonstrates a binary addition tree built from numbers and nested operations:

#include <variant>
#include <memory>
#include <iostream>

struct Node;
struct Number { double value; };
struct BinaryOp {
    char op;
    std::unique_ptr<Node> left;
    std::unique_ptr<Node> right;
};
struct Node {
    std::variant<Number, BinaryOp> expr;
    double eval() const {
        struct Visitor {
            double operator()(const Number& n) const { return n.value; }
            double operator()(const BinaryOp& b) const {
                double l = b.left->eval();
                double r = b.right->eval();
                return b.op == '+' ? l + r : 0.0;
            }
        };
        return std::visit(Visitor{}, expr);
    }
};

int main() {
    // Build (1 + 2) + 3 = 6
    auto leaf1 = std::make_unique<Node>(Node{Number{1}});
    auto leaf2 = std::make_unique<Node>(Node{Number{2}});
    auto inner = std::make_unique<Node>(Node{BinaryOp{'+', std::move(leaf1), std::move(leaf2)}});
    auto root  = std::make_unique<Node>(Node{BinaryOp{'+', std::move(inner), std::make_unique<Node>(Node{Number{3}})}});
    std::println("result: {}", root->eval());
}

std::visit dispatches to the appropriate overload based on the active alternative. The helper Visitor aggregates the lambdas (a pattern often called overloaded) and makes the code concise. Compared with a traditional C union, std::variant carries a tag, eliminating the undefined behaviour that arises from reading the wrong member. Rust’s enum offers a similar safety guarantee. The C++ version fits naturally into the existing type system without requiring a separate language construct.

The expression tree is a Prolog term in C++ clothing. A Node is either a Number (a leaf) or a BinaryOp (a compound term with two subterms). The eval function is the interpreter: it walks the term and reduces it to a value. The visit call is the case dispatch, and the Visitor struct is the set of clauses, one per alternative. This is the pattern the book returns to in chapter 22, where a compile-time SQL parser builds a similar term and evaluates it at translation time.

The full expression-tree example is a complete program. The test verifies that the tree evaluates to the correct value.

#include <variant>
#include <memory>
#include <iostream>

struct Node;
struct Number { double value; };
struct BinaryOp {
    char op;
    std::unique_ptr<Node> left;
    std::unique_ptr<Node> right;
};
struct Node {
    std::variant<Number, BinaryOp> expr;
    double eval() const {
        struct Visitor {
            double operator()(const Number& n) const { return n.value; }
            double operator()(const BinaryOp& b) const {
                double l = b.left->eval();
                double r = b.right->eval();
                return b.op == '+' ? l + r : 0.0;
            }
        };
        return std::visit(Visitor{}, expr);
    }
};

int main() {
    // Build (1 + 2) + 3 = 6
    auto leaf1 = std::make_unique<Node>(Node{Number{1}});
    auto leaf2 = std::make_unique<Node>(Node{Number{2}});
    auto inner = std::make_unique<Node>(Node{BinaryOp{'+', std::move(leaf1), std::move(leaf2)}});
    auto root  = std::make_unique<Node>(Node{BinaryOp{'+', std::move(inner), std::make_unique<Node>(Node{Number{3}})}});
    std::println("result: {}", root->eval());
}

std::optional for the maybe

A function that can return a value or nothing uses std::optional<T>. An optional either holds a T or is empty. It replaces the raw-pointer-returns-null idiom, which a caller can forget to check. The value() accessor throws std::bad_optional_access if the object is empty, and operator* is undefined for an empty optional, so the type makes the empty state explicit in the API. This pattern aligns with the Core Guidelines principle that error-prone pointer use must be avoided (F.4).

An optional carries no information about why the value is absent. It answers the question “is there a value?” and nothing more. When the caller needs to know why the operation failed, std::expected is the right type, because it carries an error description alongside the absence.

The choice between optional and expected is a question of what the caller must know. Use optional when the absence is a normal state, not an error. A map lookup that finds no key is a good fit, because the caller can proceed without the value. Use expected when the absence is a failure that the caller must handle or report. A file read that fails needs a reason, so the caller can log it or retry. The two types share the same shape, yet they carry different meaning. The book states the rule plainly: absence without a reason is optional, absence with a reason is expected.

std::expected for fallible computations

When a computation can fail with a recoverable error, std::expected<T, E> conveys either a successful result of type T or an error description of type E. The parse-integer example returns an int on success or a parse_error struct describing the position of the first non-digit character.

#include <expected>
struct parse_error { int position; std::string_view message; };
std::expected<int, parse_error> parse_int(std::string_view sv);

The caller inspects the returned object with if (result) and branches on success or failure. This approach is preferable to exceptions for anticipated failure such as malformed user input, because the control flow is explicit and does not incur the cost of stack unwinding. It also improves on std::optional by providing diagnostic information: the error type can carry a message, an error code, or any richer context the application needs. Chapter 9 expands on the trade-offs between exceptions, expected, and optional.

The parse-integer example is a complete program. The test verifies that a valid input produces the correct integer and an invalid input produces the error message.

The expected type also supports composition. A caller can chain operations so that a failure stops the chain early. The member function and_then applies a follow-up function only when the object holds a value. The member function or_else supplies a fallback when the object holds an error. These operations keep the control flow in the value domain instead of in a set of nested checks. The result reads like a pipeline of steps, and the error path stays visible in the type. This is the same spirit as the visit dispatch in the expression tree: the shape of the type drives the flow of the code.

The error type E is a full type, not a fixed string. A program can define an error that carries a code and a message, or a nested error from a lower layer. The caller can inspect the error and decide how to recover. This flexibility is what separates expected from a bare bool return. A bool tells the caller that something failed, while an expected tells the caller what failed and where. The extra information is the reason the book reaches for expected over a status flag.

#include <print>
#include <string_view>
#include <expected>

struct parse_error {
    int position;
    std::string_view message;
};

std::expected<int, parse_error> parse_int(std::string_view sv) {
    int result = 0;
    int pos = 0;
    for (char c : sv) {
        if (c >= '0' && c <= '9') {
            result = result * 10 + (c - '0');
            ++pos;
        } else {
            return std::unexpected(parse_error{pos, "non-digit character"});
        }
    }
    return result;
}

int main() {
    auto good = parse_int("123");
    if (good) {
        std::println("value: {}", *good);
    } else {
        std::println("error at {}", good.error().position);
    }
    auto bad = parse_int("12a3");
    if (bad) {
        std::println("value: {}", *bad);
    } else {
        std::println("error at {}", bad.error().position);
    }
    return 0;
}

Try this

Extend the point example from the product-type section by adding a defaulted three-way comparison (operator<=>). Create three point objects, store them in a std::vector<point>, and call std::sort. State in one sentence which comparison the sort algorithm used.

Control flow, modernly

This chapter examines modern control-flow constructs introduced in C++23 and C++26 that let the programmer keep the code that establishes a value adjacent to the test that uses it. We look at init-statements for if, switch, while and for, the ability to place structured bindings directly in a condition, the range-for loop as the preferred iteration form, compile-time conditionals with if constexpr, and the recommendation to replace manual loops with standard algorithms. By using these features the intent of the code stays close to the point of use, reducing accidental reuse and improving readability.

Init-statements

C++23 introduced init‑statements so a temporary variable can be created, tested, and destroyed within the condition of if, while, or for. A simple call yields compact syntax. For more complex setup a helper returning an RAII object can be used directly.

Init‑statements also appear in while and range‑based for. In a while loop the initializer runs once before the first test, allowing resource acquisition followed by exhaustion testing. In a for loop the initializer can bind a temporary range object, ensuring the range lives exactly for the loop’s duration. This keeps lifetimes tightly scoped and prevents accidental reuse outside the loop body.

Example: init‑statement in an if

if (auto line = std::getline(std::cin, s); !line.empty()) {
    std::cout << "first line: " << s << '\n';
}

The variable line exists only while the condition is evaluated and the body runs. No other part of the function can mistakenly read it.

This tight scoping prevents accidental reuse of the temporary buffer later in the function, which is a common source of bugs in legacy code. By limiting the lifetime, the compiler can also apply stack‑slot reuse optimizations, reducing memory pressure in tight loops.

Structured bindings in a condition (C++26)

C++26 extends init-statements by enabling structured bindings directly in the condition. A binding can decompose a tuple-like object and the boolean test can refer to any of the bound names:

if (auto [ok, n] = try_parse(s); ok) {
    std::cout << "got " << n << '\n';
}

try_parse returns a std::pair<bool,int> (or a std::expected<int,std::string_view>). The binding extracts the success flag ok and the parsed value n. The subsequent test ok decides whether the body runs. Before C++26 the language only permitted a single simple declaration inside the condition, so a common pattern was a nested if pyramid:

auto result = try_parse(s);
if (result.ok) {
    int n = result.value;
    // …
}

The new form collapses that pyramid into one line, reducing visual noise and keeping the “parse-then-use” logic together. It is the modern replacement for the nested-if pattern often used with std::expected or std::pair.

Using structured bindings in a condition also makes error handling more direct. When a function returns a std::expected, the success flag and the value can be examined immediately. This enables the error path to be written without an extra temporary variable. This leads to code that reads like a natural language description of the operation, which matches the goal of modern C++ to be expressive and intent-revealing.

#include <iostream>
#include <utility>
#include <string_view>

// Simple parser that returns {true,42} for the literal "42",
// otherwise returns {false,0}.
std::pair<bool,int> try_parse(std::string_view s) {
    if (s == "42") return {true, 42};
    return {false, 0};
}

int main() {
    std::string_view input = "42";
    if (auto [ok, n] = try_parse(input); ok) {
        std::cout << "got " << n << '\n';
    }
    return 0;
}

switch need not switch on an integer

Although the classic switch works well with integral types, modern C++ encourages the use of std::visit for variant-like data. When a developer needs to dispatch based on a value that is not an integer, the visit pattern provides exhaustive handling and integrates with concepts for compile‑time checks.

In practice, a switch on an enum can still be useful when the set of cases is closed and the compiler can emit a jump table. However, for open‑ended sets such as std::variant the visit approach avoids the risk of missing a case and produces clearer error messages.

When performance is critical, a switch with contiguous case values can be faster than a series of if‑else checks because the compiler can generate a direct table lookup. Yet, readability and safety often outweigh micro‑optimisations, especially in high‑level code.

The guidelines suggest preferring visit when dealing with sum types and reserving switch for simple, closed enumerations where the intent is to map each constant to a distinct branch.

std::variant<int,double,std::string> v = 3.14;
std::visit([](auto&& arg){
    std::cout << arg << '\n';
}, v);

For range‑based algorithms replace a switch that branches on element values with a standard algorithm such as std::count_if or std::transform. The algorithm expresses what to compute, not how to iterate.

Ranges-for is the only loop you write

A range-for iterates directly over the elements of a view or container:

for (auto&& x : container) {
    // use x
}

It eliminates the manual index variable, the off-by-one risk, and the need to look up container[i]. The Core Guidelines (ES.71) state: write a range-for whenever you need to touch each element. Index-based loops belong only to cases where the index itself is a required output.

The FizzBuzz program below runs over the view iota(1,21), which generates the integers 1 through 20. No explicit index variable appears because the view owns the counting.

#include <iostream>
#include <ranges>

int main() {
    for (int i : std::views::iota(1, 21)) {
        if (i % 15 == 0) std::cout << "fizzbuzz";
        else if (i % 3 == 0) std::cout << "fizz";
        else if (i % 5 == 0) std::cout << "buzz";
        else std::cout << i;
        if (i != 20) std::cout << ' ';
    }
    std::cout << '\n';
    return 0;
}

if constexpr

C++17 introduced if constexpr, a compile‑time conditional. The false branch is discarded before instantiation, so it need not compile. This lets generic code adapt to template arguments via constant‑expression predicates such as std::is_integral_v<T>. The selected branch can be inlined, eliminating dead code and improving optimisation.

A typical use tests std::ranges::range<T> to choose between printing a scalar value or iterating a range. Concepts can be combined, e.g. requires { typename T::value_type; }, to keep constraints close to the code they govern.

Overall, if constexpr provides clear, type‑safe compile‑time dispatch without separate specialisations.

Raw loops are a smell

When an algorithm expresses the intent (e.g., find the first element greater than 10), writing a manual loop hides that intent. The guideline (ES.70) urges the programmer to replace a hand‑written loop with an appropriate standard algorithm, such as std::find_if, std::count, or std::transform, so that readers instantly recognise the operation. Chapter 12 will enumerate the algorithm zoo and show how each algorithm maps to a common pattern.

Standard algorithms also bring strong exception‑safety guarantees. Because the library implements the iteration and cleanup logic, edge cases such as early exits or thrown exceptions are handled uniformly. This reduces the likelihood of resource leaks compared with manually written loops that must explicitly manage cleanup. Moreover, many algorithms are annotated for vectorisation. The compiler can then generate SIMD instructions automatically when the iterator type supports it. The result is often faster code with the same expressive clarity.

When performance is critical, the programmer can still profile the algorithmic choice. The standard library provides overloads that accept execution policies. These enable parallel execution without changing the high‑level code. This flexibility means that the same source can be tuned for different hardware targets by swapping the policy argument.

Overall, preferring algorithms over hand‑written loops aligns with modern C++ philosophy: write what you want to achieve, let the library handle how to achieve it.

Try this

Rewrite the FizzBuzz program so that the upper bound is supplied by a function parameter and the program prints only the count of numbers that are multiples of both 3 and 5. Use the std::views::filter adaptor to select the qualifying numbers and std::ranges::distance to obtain the count.

Ownership, move, and RAII

Value semantics as the default mental model

C++ treats a variable as the sole owner of the value it stores. The compiler creates the value when the variable’s defined and destroys it when the variable’s lifetime ends. This model matches the Rust model without the borrow checker: a name owns a value.

Passing the name to a function transfers the value or copies it. The transfer occurs by invoking a move constructor or a copy constructor. The compiler later checks that every move obeys the lifetime‑safety rules introduced in Chapter 1.

Contrast this model with raw C pointers. A pointer refers to memory that any code path can own. The language does not track that ownership. Programmers conventionally treat the pointer as borrowed, but the compiler cannot enforce that rule. Consequently, dangling pointers and double frees appear frequently in legacy code.

In modern C++ code, the default assumption is ownership per value. When a function receives a parameter by value, the caller gives up its ownership. When a function returns a value, the caller receives a fresh owner. The only time a value is duplicated is when a copy constructor executes.

The contrast with Rust is direct. Rust’s borrow checker proves at compile time that no two owners exist for the same value and that no reference outlives its referent. C++ has no borrow checker, so the same invariants rest on type design: owning types delete copy, non-owning views borrow without duplicating, and the lifetime analysis flags dangling references. The mental model is the same. The enforcement differs.

Special member functions and the Rule of Zero/Five

A type can define up to six special member functions: default constructor, destructor, copy constructor, move constructor, copy-assignment operator, and move-assignment operator. If a class contains only owning standard‑library members (e.g., std::vector, std::string, std::unique_ptr), the compiler‑generated versions are correct. This is the Rule of Zero: write no special members and rely on the compiler.

If a class manages a raw resource unknown to the language (e.g., a C FILE*), it must provide a destructor, delete copy operations, and implement a move constructor that transfers ownership. This follows the Rule of Five. The Rule of Zero is the default, cheapest correct choice. The Rule of Five applies only at system boundaries where a raw resource is wrapped.

Value categories: lvalues and rvalues

An expression that has a name and an addressable location is an lvalue. An unnamed temporary or a cast that produces a temporary is an rvalue. An rvalue is about to be destroyed, and the compiler can reuse its resources.

C++ refines rvalues into two subcategories. A prvalue is a pure temporary, the result of an expression like 42 or std::string("hi"). An xvalue is an expiring value, an object that has a name but is about to be destroyed, which is what std::move produces. Both are rvalues. The full ladder:

CategoryHas a name?Can be moved from?Example
lvalueyesnoint x; then x
prvaluenoyes42, std::string("hi")
xvalueyesyesstd::move(x)
rvalue(prvalue or xvalue)yes(either of the above)

The distinction is not about what the expression is, but about what the program will do with it next. A named variable is an lvalue because the program will read it again. A temporary is an rvalue because nobody will read it again after the current expression. The move constructor is the mechanism that lets the compiler steal from the rvalue instead of copying it.

std::move is a cast that converts an lvalue to an rvalue. It does not move any bytes. It merely tells the compiler that the object can be treated as a disposable source. If a move constructor exists, the compiler selects it. Otherwise it falls back to the copy constructor.

Because std::move is a cast, you can apply it repeatedly without side effects. The actual movement happens exactly once, when the move constructor runs.

The two categories matter because they decide what the compiler is allowed to do with the object. An lvalue has a name and a future: the program will use it again, so the compiler must keep its value intact. An rvalue is about to be destroyed: nobody will read it again, so the compiler is free to steal its resources. std::move is how you tell the compiler that an lvalue is now in the second category. The cast is safe only if you will not use the source again after the move, because the source is left in a valid-but-unspecified state.

Move is cheap. copy is honest

Moving a std::vector<int> copies three pointers: the data address, the size, and the capacity. The source vector becomes empty. Copying a std::vector<int> allocates new storage and copies every element. The cost difference is usually orders of magnitude.

The same ratio holds for std::string, std::list, and every container that owns heap memory. A move is a pointer swap. A copy is an allocation plus a loop. This is why the book treats move as the default transfer mechanism and copy as the deliberate, expensive choice.

Standard containers require their move constructors to be noexcept (CG C.66). When a container reallocates during growth, it prefers to move elements rather than copy them. If the move constructor can throw, the container falls back to copying to preserve strong exception safety. Therefore you must mark move constructors noexcept whenever possible.

The noexcept guarantee is a contract between your type and the container. A type whose move constructor throws breaks the strong exception guarantee when the container grows, because a throw mid-reallocation leaves the container in a state where some elements moved and some did not. The container refuses to take that risk: it copies instead, which is correct but slow. Marking the move noexcept is the one-word fix that lets the container take the fast path.

RAII as the central idiom

Resource Acquisition Is Initialization (RAII) states that a resource is obtained in a constructor and released in the matching destructor (CG R.1, R.3). The file_guard struct illustrates this pattern: the constructor calls std::fopen. The destructor calls std::fclose.

Contrast the RAII version with manual management:

std::FILE* f = std::fopen("tmp.txt","w");
std::fwrite("X",1,1,f);
std::fclose(f);

If the function returns early, the explicit std::fclose is skipped and the file leaks. The RAII wrapper eliminates that risk because the destructor runs automatically on every exit path, including exceptions.

RAII wrapper for a C file handle

The following example compiles the target ch05_file_guard. The test expects the word closed on standard output, confirming that the destructor runs when the scope ends.

#include <cstdio>
#include <print>

struct file_guard {
    std::FILE* f_ = nullptr;
    explicit file_guard(const char* path, const char* mode) : f_(std::fopen(path, mode)) {
        if (f_) std::println("opened");
    }
    // Delete copy operations – a file handle cannot be duplicated safely.
    file_guard(const file_guard&) = delete;
    file_guard& operator=(const file_guard&) = delete;
    // Move transfers ownership; source is left null.
    file_guard(file_guard&& other) noexcept : f_(other.f_) {
        other.f_ = nullptr;
        if (f_) std::println("moved");
    }
    file_guard& operator=(file_guard&& other) noexcept {
        if (this != &other) {
            if (f_) std::fclose(f_);
            f_ = other.f_;
            other.f_ = nullptr;
            if (f_) std::println("move‑assigned");
        }
        return *this;
    }
    ~file_guard() {
        if (f_) {
            std::fclose(f_);
            std::println("closed");
        }
    }
    void write_byte(char c) {
        if (f_) std::fputc(c, f_);
    }
};

int main() {
    {
        file_guard fg("tmp.txt", "w");
        fg.write_byte('A');
    }
    return 0;
}

Instrumented move and copy

The program prints copy when the line blob b = a; executes. The std::move call triggers the move constructor, which prints move. The test expects the word copy, proving that the first operation is a copy.

#include <vector>
#include <print>

struct blob {
    std::vector<int> data;
    // Default constructor
    blob() = default;
    // Copy constructor – prints "copy"
    blob(const blob& other) : data(other.data) {
        std::println("copy");
    }
    // Move constructor – prints "move"
    blob(blob&& other) noexcept : data(std::move(other.data)) {
        std::println("move");
    }
    // Copy assignment – prints "copy assign"
    blob& operator=(const blob& other) {
        data = other.data;
        std::println("copy assign");
        return *this;
    }
    // Move assignment – prints "move assign"
    blob& operator=(blob&& other) noexcept {
        data = std::move(other.data);
        std::println("move assign");
        return *this;
    }
    // Destructor – defaulted; data cleans itself up.
    ~blob() = default;
};

int main() {
    blob a;                // default constructed
    a.data = {1,2,3};
    blob b = a;            // copy – should print "copy"
    blob c = std::move(a); // move – should print "move"
    (void)b; (void)c;      // silence unused warnings
    return 0;
}

In this book, every resource lives inside an object whose destructor frees it. You never see a raw fopen/fclose pair outside a RAII wrapper.

RAII is the bridge between the C world of manual resource management and the C++ world of value semantics. The wrapper is the only place that touches the raw API. Everything else sees a type that moves, destructs, and cleans up automatically. This is why the chapter shows new and delete exactly once and then drops them: the raw primitives exist to explain what the wrapper replaces, not to be used directly.

new and delete appear only once, then disappear

The classic C-style allocation pattern looks like this:

int* p = new int{42};
delete p;

That pattern appears only in the explanatory paragraph above. All other code relies on standard library containers, std::unique_ptr, or RAII wrappers. The guideline is never write new or delete in production code. Let the compiler generate those calls inside containers or smart pointers.

The reason is not aesthetic. A raw new is a leak waiting to happen: every early return, every exception, and every forgotten delete leaks the object. A container or smart pointer ties the deallocation to a destructor, which the compiler guarantees to run. The two lines above are the only place in the book where the raw primitives appear, and they exist to make this argument concrete.

Try this

Write a struct blob that deletes its copy constructor and copy-assignment operator, but keeps the move constructor and move-assignment operator. Then call a function void consume(blob b); with std::move on a local blob variable. Verify that the call compiles, while a plain copy blob c = b; fails to compile. The exercise demonstrates that moving a value does not require a copy and that the compiler enforces the deletion.

Smart pointers and owning views

Unique ownership is the default

The Core Guidelines (R.20, R.21) require that a heap object have exactly one owning smart pointer. std::unique_ptr<T> satisfies this rule. It holds a pointer, destroys the object when the unique_ptr itself is destroyed, and can be moved but never copied. The move operation transfers the stored pointer and leaves the source empty.

#include <memory>
#include <vector>
#include <print>

struct tree_node {
    int value{};
    std::vector<std::unique_ptr<tree_node>> children;
};

void print(const tree_node& node) {
    std::print("{} ", node.value);
    for (const auto& child : node.children) {
        print(*child);
    }
}

int main() {
    auto root = std::make_unique<tree_node>();
    root->value = 1;
    tree_node* cur = root.get();
    for (int v = 2; v <= 5; ++v) {
        cur->children.emplace_back(std::make_unique<tree_node>());
        cur->children.back()->value = v;
        cur = cur->children.back().get();
    }
    print(*root);
    std::println("");
    return 0;
}

In the tree example the root is a std::unique_ptr<tree_node>. Each node stores its children in a std::vector<std::unique_ptr<tree_node>>. Because every child is owned uniquely, the destruction of the root automatically destroys the whole subtree recursively. No manual delete appears, and the program cannot accidentally copy a node-owner. The compiler rejects any copy of a unique_ptr.

Transfer of ownership by value

A function that receives a unique_ptr<T> by value takes ownership. The caller must move the pointer into the parameter. After the call the caller’s pointer becomes empty. This pattern appears in many factory functions:

std::unique_ptr<tree_node> make_root(int v) {
    auto p = std::make_unique<tree_node>();
    p->value = v;
    return p; // move-return, caller receives ownership
}

In the tree example the statement auto root = std::make_unique<tree_node>() creates the sole owner. When main ends, root goes out of scope, the move-return chain unwinds, and the destructor of each unique_ptr in the vectors frees the corresponding child. No memory leaks survive past main.

std::make_unique versus new

The guidelines (R.23) require that a unique_ptr be created with std::make_unique rather than a raw new expression. std::make_unique<T>(args...) constructs the object and wraps it in a unique_ptr in one step. This form is shorter and safer than the two-step alternative. The two-step form first calls new T(args...) and then passes the result to the unique_ptr constructor. If an exception occurs between these two steps, the raw pointer leaks. std::make_unique avoids this window, because the construction and the wrapping happen together. The function also deduces the type, so the code does not repeat the type name. Use std::make_unique whenever the object is created and owned immediately. Reserve a raw new for the rare case where a custom deleter or a pre-existing pointer is required.

Move semantics give unique_ptr its efficiency. Moving a unique_ptr transfers the pointer without copying the pointed-to object. The operation is a simple pointer assignment plus a null-out of the source. It never allocates and never touches the heap object. This property lets a function return a unique_ptr by value at no cost. The move-return in make_root does not copy the tree. It hands the same object to the caller. The same reasoning applies when a unique_ptr moves into a container or into another owner.

Shared ownership is a cost-aware choice

std::shared_ptr<T> holds a control block with an atomic reference count. Every copy increments the counter. The last copy decrements to zero and destroys the object. The guidelines (R.22) caution that shared_ptr must be used only when at least two distinct owners truly need to keep the object alive. The cost model is higher than unique_ptr:

Aspectunique_ptrshared_ptr
Allocationone block (object)two blocks (object + control)
Reference countnoneatomic increment/decrement on every copy
Size of pointer object≤ sizeof(void*)≈ 2 × sizeof(void*)
Cache behaviorcontiguous accessindirect control block
If the program never needs more than one owner, unique_ptr is the zero‑overhead choice. shared_ptr allocates an extra control block and incurs atomic reference‑count updates on each copy, doubling pointer size and reducing cache locality. Use shared_ptr only when at least two owners truly need to keep the object alive. Otherwise the extra cost is unnecessary.

Weak pointers break cycles

std::weak_ptr<T> provides a non‑owning view that does not affect the reference count, breaking reference cycles. Use lock() to obtain a temporary shared_ptr if the object is still alive. Otherwise lock() returns an empty pointer. expired() reports whether the object has been destroyed. In the graph example, weak_ptr allows the parent link to be observed without extending the child’s lifetime, preventing a cycle.

#include <memory>
#include <vector>
#include <print>

struct graph_node {
    int id{};
    std::weak_ptr<graph_node> parent;
    std::vector<std::shared_ptr<graph_node>> children;
};

void print_ids(const graph_node& node) {
    std::print("{} ", node.id);
    for (const auto& child : node.children) {
        print_ids(*child);
    }
}

int main() {
    auto a = std::make_shared<graph_node>();
    a->id = 1;
    auto b = std::make_shared<graph_node>();
    b->id = 2;
    // create edge a -> b and back-edge parent weak_ptr
    a->children.push_back(b);
    b->parent = a; // weak, no cycle
    print_ids(*a);
    std::println("");
    // when main exits, both a and b are reclaimed; no leak.
    return 0;
}

gsl::owner<T> documents raw owning pointers

Sometimes a C-language API requires a raw pointer that owns the pointed-to object. The Guidelines Support Library provides gsl::owner<T> as a type alias that makes the ownership intent explicit to static analysis tools:

void c_api(gsl::owner<int*> p); // p must be freed by the caller

gsl::owner<T> is a type alias that documents raw owning pointers for static analysis tools such as Clang’s lifetime‑safety analysis. It carries no runtime cost and does not alter program behavior. A function taking a gsl::owner<T> parameter signals that it assumes ownership. A function returning gsl::owner<T> signals that it transfers ownership to the caller. Static analysis can verify that each owner releases the object exactly once.

When pointers are the wrong tool

Ownership and lifetime are not the only reasons to use a pointer. Frequently a value, a std::span, or a std::string_view conveys the required relationship without any ownership semantics.

  • Value: use when the object’s lifetime is confined to the current scope and copying is cheap.
  • std::span<T>: a non-owning view over a contiguous range. It is ideal for passing array slices to functions.
  • std::string_view: a read-only view of a string. It is perfect for read-only parameters where the callee must not modify or own the data. Choosing a smart pointer when a simple view suffices adds unnecessary indirection and can hide bugs. The guidelines (R.3) advise to prefer plain values and views first. Smart pointers come into play only when the lifetime must outlive the current scope and no value can express the relation. A value owns its data within the current scope. Copying or moving it transfers ownership automatically. std::span<T> provides a non‑owning view of a contiguous range, and std::string_view provides a read‑only view of a string. These types express the relationship without ownership overhead and avoid the need for raw pointers. Use them before considering a smart pointer.

Try this

Take the tree program from the previous section and add a function:

int height(const tree_node&);

height returns the length of the longest root-to-leaf path (the number of nodes on that path). Write the function recursively, using only the tree_node interface. When you run the program, observe that the tree is still freed automatically when main exits, even though height returns a plain int. What does std::unique_ptr guarantee about the tree’s memory after height returns?

Lifetimes, and the compiler that sees them

Storage durations in one breath

C++ classifies every object by its storage duration, which determines when memory is allocated and reclaimed. The three categories cover nearly all cases:

  • Automatic objects are created when control enters their block and destroyed on exit. Their lifetime matches the block’s scope.
  • Static objects exist for the program’s lifetime, created before main and destroyed after it returns.
  • Dynamic objects live on the heap, managed by containers or smart pointers that acquire and release the memory.

The contrast with C is sharp. In C, malloc allocates heap memory without an automatic scope link. The programmer must call free at the exact moment, separating pointer and memory lifetimes. C++ eliminates this gap: containers or smart pointers own heap objects, and those owners have automatic or static lifetimes, so reasoning focuses on scope and ownership rather than matching allocation and deallocation.

RAII links storage duration to resource lifetime. An automatic object that owns a resource releases it in its destructor when the block ends, so the resource’s lifetime matches the object’s lifetime. This enables the compiler to understand and enforce lifetimes.

The dangling taxonomy

The Core Guidelines Lifetime profile enumerates four ways a program can keep using a value after the value’s storage is gone. Each case has a fixed shape, and each can be caught by static analysis when the analysis runs.

  1. Return of a local reference or pointer. A function returns the address of a variable that goes out of scope when the function returns. The returned handle points at memory the compiler is about to reuse.
  2. Use after end of scope or delete. Code reaches through a handle after the object’s destructor has run, or after delete has freed the memory. The handle is live but the value is dead.
  3. Views that outlive their source. A std::string_view, an iterator, or a raw pointer keeps being used after the object it observes dies or moves. The view borrows storage that no longer holds the expected value.
  4. Container invalidation. Inserting or erasing elements can trigger reallocation, which moves the underlying storage. Any pointer, reference, or iterator taken before the operation now points at freed or stale memory.

These four cases are not arbitrary. They are the only ways a handle can survive its referent, because a handle can only outlive its referent through a return, through a delayed use, through a non-owning view, or through a container resize. The compiler can model each shape because each shape is a small, local pattern in the code.

Lifetime-safety analysis as a compiler feature

The analysis is the Clang implementation of the Core Guidelines Lifetime Safety profile. The profile was written to make the rules above checkable, and the Clang work turned those rules into a dataflow pass over the abstract syntax tree. The stable flag is -Wlifetime-safety. On the pinned toolchain, Clang 22.1.8, that stable flag does not exist yet. The experimental form -Xclang -fexperimental-lifetime-safety is available, and it emits only the -Wreturn-stack-address warning. The deeper checks for cases two, three, and four remain silent in this build.

The book treats the analysis as a law the compiler will enforce when the feature matures. The same committee member who advanced the flag in Chapter 01 is behind the wider effort. The point of teaching the taxonomy now is that the discipline is correct regardless of whether the tool shouts at you. A program that obeys the lifetime rules is correct on today’s compiler and will stay correct on tomorrow’s.

Demo: return of a stack address

// warning: address of stack memory associated with local variable 'x' returned [-Wreturn-stack-address]
#include <cstddef>

int* dangling_pointer() {
    int x = 42;
    return &x; // returns address of a local variable
}

int main() {
    int* p = dangling_pointer();
    // Using *p here would be undefined behavior
    (void)p;
    return 0;
}

The warning at the top of the file is the diagnostic Clang emits when a function returns a reference to a local variable. The local lives in the function’s stack frame. When the function returns, that frame is reclaimed, and the returned reference points at storage the next call is free to overwrite. No runtime check can save this case. The value is gone before anyone reads it. The fix is to return by value, or to have the caller pass the destination in and let the function fill it through a reference the caller owns.

Demo: vector invalidation

#include <vector>
#include <iostream>

int main() {
    std::vector<int> v{1, 2, 3};
    int* p = &v[0]; // capture pointer to first element
    v.push_back(4); // can cause reallocation, invalidating p
    // If compiled with AddressSanitizer, a heap-use-after-free error would be reported when *p is accessed.
    std::cout << *p << "\n"; // undefined behavior if reallocation occurred
    return 0;
}

The program captures a pointer to the first element, then calls push_back. A std::vector keeps its elements in a contiguous block. When that block is full, push_back allocates a larger block and moves every element into it. The old block is freed. The captured pointer still points at the old block, so reading through it is a use of freed memory. When compiled with AddressSanitizer the runtime reports heap-use-after-free.

The lesson is general. Any operation that can grow a std::vector can invalidate pointers, references, and iterators into it. The safe pattern is to take the address after the last growth, to call reserve up front when the final size is known, or to keep using the vector’s own indexing instead of a captured handle.

Demo: dangling string_view

#include <string_view>
#include <string>

std::string_view dangling_view() {
    std::string temp = "temporary"; // lives only inside function
    return std::string_view{temp}; // view refers to destroyed string
}

int main() {
    auto sv = dangling_view();
    // Using sv here is undefined behavior because the source string has been destroyed.
    (void)sv;
    return 0;
}

The function returns a std::string_view that refers to a temporary std::string. The temporary is destroyed at the end of the full expression that created it, while the view is returned to the caller. The caller then reads through a view whose characters are already gone. This is undefined behavior, and it is silent. The code compiles and can even appear to work until the memory is reused.

The trap is that std::string_view is cheap to return, so it invites returning a view into a value the function just made. The rule is to return a std::string when the source is a temporary, and to return a std::string_view only when the caller already owns the backing storage for the whole time the view is used.

Demo: lifetime‑bound annotation

#include <string_view>
#include <string>

// First version: no lifetime annotation. The analysis does not warn when a temporary is passed.
std::string_view first_word(const std::string& s) {
    // Return view to first word (up to first space)
    auto pos = s.find(' ');
    if (pos == std::string::npos) return std::string_view{s};
    return std::string_view{s.data(), pos};
}

int main() {
    // Passing a temporary string – the view dangles, but current clang gives no diagnostic.
    auto sv = first_word(std::string{"hello world"});
    (void)sv;
    return 0;
}

/*
// Fixed version with lifetimebound annotation (Clang accepts the attribute).
std::string_view first_word(const std::string& s) [[clang::lifetimebound]] {
    auto pos = s.find(' ');
    if (pos == std::string::npos) return std::string_view{s};
    return std::string_view{s.data(), pos};
}
*/

The first version returns a view into its parameter without any annotation, so the compiler does not warn when a temporary is passed. Adding [[clang::lifetimebound]] to the parameter tells the analysis that the returned view is tied to the argument’s lifetime. Current Clang accepts the attribute but emits no diagnostic. The programmer must still respect the contract, which future compilers will enforce.

The attribute can also be placed on the function itself, applying the rule to every return path, and on constructors: a constructor that stores a pointer or reference member must mark the source parameter lifetimebound so the member’s lifetime cannot outlive the argument. This links the member to the argument and enables static checks.

The pattern appears wherever a function hands back a handle into memory it was given. Accessors that return a reference or a view to a member, parsers that return a view into their input, and span factories all need the attribute to connect output to input.

Views and spans are lifetime-transparent

std::string_view and std::span are non-owning handles. They do not own storage. They merely borrow it. A view has no destructor that frees anything, because there is nothing for it to free. Its correctness depends entirely on the owner staying alive.

Keep the owner alive for at least as long as any view or span that refers to it. A view is a loan. The lender must outlive the loan. Creating the owner and view in the same expression makes the temporary owner die before the view is used. Bind the view to a named owner that lives in an enclosing scope to avoid the hazard.

std::span generalizes the idea from characters to any T. A std::span<const int> is a non-owning window over a row of int values. The same lifetime rule applies. The std::vector<int> or array that backs the span must outlive every use of the span. Because a span is so light, functions must accept spans instead of a raw pointer plus a length. That change makes the non-owning intent explicit and removes a whole class of length mismatch bugs.

gsl::not_null<T>

The guideline support library provides gsl::not_null<T>. It wraps a raw pointer and guarantees that the pointer is never null. The wrapper adds a contract that the analysis can check. In practice, references already enforce non-nullness, so the book uses gsl::not_null only when a raw pointer is required for legacy interop. The lifetime rules still apply on top of the nullness rule. A gsl::not_null that points at a destroyed object is just as dangling as a plain pointer.

The two contracts are independent. gsl::not_null answers the question of whether the pointer is null, while the lifetime rules answer the question of whether the referent is still alive. A handle can satisfy both and still dangle, so the wrapper does not remove the need for the lifetime discipline taught in this chapter.

Try this

Write a function std::span<const int> middle(std::vector<int>& v) that returns a span over the elements v[1] through v[n-2]. State the rule that the caller must keep v alive for the span’s lifetime. Explain what happens if v is a temporary, and show the [[clang::lifetimebound]] marking that makes the contract explicit.

Passing arguments

Pass by value vs by reference

In C++ a function parameter can be a value, a reference, or a pointer. The choice determines whether the callee receives its own copy of the argument or an alias to the caller’s object.

ParameterTypical useEffect on caller
T (by value)The function needs its own mutable objectThe argument is copied or moved. The caller’s object is unchanged
const T&The function only reads the argumentNo copy is performed. The callee cannot modify the object
T&The function must modify the caller’s objectThe caller sees the mutation. No copy is performed
T*Legacy C‑style API or optional argumentThe pointer can be null. The callee can reassign the pointer

The rule of thumb is:

  • Pass by value when the function will create a new object, store it, or move from it.
  • Pass by const reference when the function only needs to observe the argument.
  • Pass by non‑const reference only when the function must change the caller’s object.

Contrast this with C. In C every argument is passed by value. To share an object the programmer writes a pointer parameter (T* p). The reader must locate * to know that the function can observe or mutate shared state. C++ hides that pointer indirection behind references. This makes the intent visible in the function signature.

Size is not the only driver. A type that is expensive or impossible to copy, such as std::unique_ptr or a large std::vector, must travel by reference or by rvalue move, never by value copy. For a value that the function only inspects, const T& is the default even when T is small, because it avoids a copy and works for every argument type. Pass by value is the right default only when the function keeps or modifies a copy of what it was given.

Arguments of differing size

#include <iostream>
#include <vector>
#include <string>

struct Small {
    int x;
    Small(int v) : x(v) { std::cout << "Small ctor\n"; }
    Small(const Small& other) : x(other.x) { std::cout << "Small copy\n"; }
    Small(Small&& other) noexcept : x(other.x) { std::cout << "Small move\n"; }
    Small& operator=(const Small&) = delete;
    Small& operator=(Small&&) = delete;
    ~Small() = default;
};

void by_value(Small s) {
    std::cout << "by_value address " << &s << "\n";
}

void by_const_ref(const Small& s) {
    std::cout << "by_const_ref address " << &s << "\n";
}

void by_mut_ref(Small& s) {
    std::cout << "by_mut_ref address " << &s << "\n";
    s.x += 1;
}

void vec_by_value(std::vector<int> v) {
    std::cout << "vec_by_value size " << v.size() << "\n";
}

void vec_by_const_ref(const std::vector<int>& v) {
    std::cout << "vec_by_const_ref size " << v.size() << "\n";
}

int main() {
    Small a(5);
    std::cout << "original address " << &a << "\n";
    by_value(a);
    by_const_ref(a);
    by_mut_ref(a);
    std::cout << "after mutation value " << a.x << "\n";

    std::vector<int> big(1000, 1);
    vec_by_value(big);
    vec_by_const_ref(big);
    return 0;
}

Running the program prints the address of each parameter. The by_value call receives a copy, therefore its address differs from the caller’s object. The by_const_ref and by_mut_ref calls receive the same address, confirming they are aliases. The vector examples show that a large container passed by value incurs a copy of its control block, while a const reference avoids any copy.

Const correctness

A const qualifier promises not to modify the object it decorates. In a function signature:

void f(const T& arg);

the callee can call only const member functions on arg. Attempting to call a non‑const member function triggers a compile‑time error because the implicit this pointer is const. The same rule applies to member functions:

struct S {
    void mutate()       { ++value; }          // non‑const
    void inspect() const { std::cout << value; } // const
    int value;
};

When inspect is called on a const S object, the compiler guarantees that value is not altered. If a const reference were bound to a temporary, the temporary becomes immutable for the duration of the reference. The compiler enforces this rule even though the temporary will soon be destroyed.

The same const placement rules apply to pointers. See Chapter 2.

If we modify ch08_pass_by.cpp to call a non‑const member on a const reference, the compilation fails, illustrating that const truly prevents mutation.

The discipline pays off via const overloading: a type can provide both T& at(std::size_t) and const T& at(std::size_t) const. The former binds to non‑const objects, and the latter binds to const objects and temporaries. This enables read‑only access such as std::string{}.size().

Pass by value then move

The “sink” idiom accepts a parameter by value and immediately moves it into a member or another container. This design gives two useful behaviors:

  • When the caller supplies an lvalue, the argument is copied into the parameter and then moved into the destination. The copy costs one construction. The move costs no additional allocation.
  • When the caller supplies an rvalue (a temporary), the argument is constructed directly in the parameter slot and then moved. This results in zero copies.

The diagram below shows both paths.

caller lvalue          caller rvalue
   |                       |
   v copy                   v construct
parameter (value)   <--->  parameter (value)
   | move                     | move
destination                 destination

Sink idiom

#include <iostream>
#include <string>

struct Small {
    int x;
    Small(int v) : x(v) { std::cout << "Small ctor\n"; }
    Small(const Small& other) : x(other.x) { std::cout << "Small copy\n"; }
    Small(Small&& other) noexcept : x(other.x) { std::cout << "Small move\n"; }
    Small& operator=(const Small&) = delete;
    Small& operator=(Small&&) = delete;
    ~Small() = default;
};

void take_and_store(Small s) {
    std::cout << "parameter address " << &s << "\n";
    Small member = std::move(s);
    std::cout << "member address " << &member << "\n";
}

int main() {
    Small a(1);
    std::cout << "original address " << &a << "\n";
    take_and_store(a);          // lvalue path – copy
    take_and_store(Small(2));   // rvalue path – move
    return 0;
}

The program prints the address of the caller’s object, the parameter, and the member after the move. When the first call uses an lvalue, a copy occurs. The second call uses a temporary, so only the move runs. The output confirms the expected behavior.

Return by value

Returning a value by copy used to be expensive because the caller received a separate object that required a copy. Modern C++ solves this with copy elision, including the guaranteed elision of prvalues, and with move semantics for named temporaries. The compiler constructs the return object directly in the caller’s storage (NRVO) or treats the temporary as an rvalue that can be moved. Consequently, returning a std::vector or a std::string does not copy the underlying buffer on the happy path.

The rule is simple: return‑by‑value is cheap because copy elision and move semantics eliminate unnecessary copies. Elision is guaranteed when the function returns a single local or a prvalue. Otherwise the compiler falls back to a move. Return a clearly identified local or construct the result directly in the return statement (see Chapter 05 for container move semantics).

References are not pointers

A reference is an alias bound to an object at initialization. It cannot be reseated and cannot be null. The language guarantees that a reference denotes a valid object for its lifetime.

The reference does not extend the lifetime of the object it refers to. If the referent is destroyed while the reference remains alive, using the reference yields undefined behavior. This hazard is described in Chapter 07.

A pointer can be null or reassigned and offers no guarantee that the pointee remains alive. Unlike Rust’s borrow checker, C++ relies on the programmer. Chapter 07’s lifetime analysis tool can detect violations. Rule: a reference must never outlive its referent.

There is one exception that surprises even experienced programmers. Binding a const reference to a temporary extends the temporary’s lifetime to match the reference’s scope. The temporary is not destroyed at the end of the full expression. It lives until the reference goes out of scope. This extension does not apply when the reference is a member of an object or when it is returned from a function. It is a property of the local binding only. It is why const auto& x = compute() avoids a copy of a temporary safely.

reference_wrapper and views as parameters

Sometimes a function must store a reference in a container or type‑erase the argument. std::reference_wrapper<T> wraps a reference in an assignable object. This allows it to appear in std::vector or std::function. The wrapper forwards operators to the underlying reference.

A std::vector<T&> is ill formed, because a reference is not a complete object that a container can own. std::reference_wrapper<T> solves this. It is a copyable, assignable object that stores a pointer to the referent and behaves like the reference in almost every context. This is the standard way to keep a collection of aliases into other objects.

For non‑owning sequences, the standard library provides view types:

  • std::string_view: a read‑only window over a character array.
  • std::span<T>: a read‑only or mutable view over a contiguous range of T.

Both types carry a pointer to the data and a length. They replace the old “pointer + size” pattern:

void process(std::span<const int> data);

The caller can pass a std::vector<int>, a C‑array, or a pointer with length. The callee receives a lightweight handle that does not own the elements. The same lifetime rules from Chapter 07 apply: the owner must outlive any view or span.

std::forward and perfect forwarding

Templates that accept universal references (T&&) can forward their argument preserving its value category. The pattern:

template<class T>
void wrapper(T&& t) {
    target(std::forward<T>(t));
}

If t is an lvalue, std::forward<T>(t) yields an lvalue reference. If t is an rvalue, it yields an rvalue reference. This enables factories, wrappers, and higher‑order functions to forward arguments without unintentionally copying or moving them.

The cost of omitting std::forward is concrete. A function that takes T&& t and then passes t to another function without forwarding always passes an lvalue, because inside the body t is a named variable. That forces a copy where the caller expected a move. std::forward<T>(t) restores the original category. An lvalue argument stays an lvalue and a temporary stays a temporary. Every standard-library helper that accepts a forwarding reference, from std::make_unique to container emplace, relies on this.

Perfect forwarding

#include <iostream>
#include <type_traits>
#include <utility>

void target(int& i) {
    std::cout << "target received lvalue reference, value " << i << "\n";
}

void target(int&& i) {
    std::cout << "target received rvalue reference, value " << i << "\n";
}

template<class T>
void wrapper(T&& t) {
    std::cout << "wrapper forwarding, is lvalue? " << std::boolalpha
              << std::is_lvalue_reference<decltype(t)>::value << "\n";
    target(std::forward<T>(t));
}

int main() {
    int x = 42;
    wrapper(x);    // lvalue path
    wrapper(100);  // rvalue path, creates temporary int
    return 0;
}

The program reports whether the forwarded argument is an lvalue. The first call forwards an lvalue (x). The second call forwards a temporary integer created from the literal 100. The target function receives the correct reference type in each case.

Try this

Write a function

std::vector<int> take_and_double(std::vector<int> v);

that doubles every element and returns v. Call it once with an lvalue and once with a prvalue. Observe (by printing the address of the parameter) that the prvalue call avoids a copy.

Errors and contracts

The C legacy

C reports failure via integer return codes and the global variable errno. Callers must remember to test each result, otherwise silent bugs arise.

C++ adds mechanisms that make ignoring failures harder: the type system can encode error states, and the [[nodiscard]] attribute warns when a result is discarded, ensuring the intent to handle errors is visible.

Exceptions and unwinding

C++ supports throw, try, and catch. A throw aborts the current function, triggers stack unwinding, and destroys each local object in reverse order, allowing RAII objects to release resources automatically and preventing leaks.

Exceptions apply when the failure cannot be repaired at the call site, such as out‑of‑memory, corrupted files, or invalid user input.

The dividing line between exceptions and expected is the frequency and the locality of the failure. A parse error is routine: the caller expects it, handles it, and moves on, so it travels as an expected value. An out-of-memory error is rare and crosses abstraction boundaries: no individual caller can fix it, so it travels as an exception and unwinds to the nearest handler. Mixing the two is a smell: if every caller wraps a function in a try/catch, the failure is routine and belongs in an expected.

Modern ABIs place the exception handling code on the cold path only. The common case executes without extra branches or hidden calls. The claim that exceptions add overhead no longer applies to the common case. When exceptions are disabled with -fno-exceptions, the compiler omits unwind tables, reducing binary size and improving load time.

noexcept

The specifier noexcept promises that a function will not throw. If a noexcept function does throw, the runtime calls std::terminate.

noexcept is a contract, not a hint. The compiler uses it to make decisions: a std::vector will move elements during reallocation only if the move constructor is noexcept, and the optimizer can elide exception-handling machinery around a noexcept call. Breaking the contract by throwing from a noexcept function calls std::terminate, which is a hard stop. The specifier is therefore a promise you make when you can guarantee it, not a wish.

Marking move constructors as noexcept enables the standard library to move objects during vector growth without a fallback copy. The optimizer can also inline the call more aggressively. noexcept functions can be inlined without hidden control flow. The optimizer gains full visibility of the call graph.

A conditional noexcept(...) expression makes the promise precise: the function is noexcept only when every operation it calls is noexcept itself, so the guarantee evolves with the implementation.

#include <iostream>
#include <utility>

struct Counter {
    int value;
    Counter(int v) : value(v) { std::cout << "constructed " << value << "\n"; }
    ~Counter() { std::cout << "destroyed " << value << "\n"; }
    // Move constructor promised not to throw
    Counter(Counter&& other) noexcept : value(other.value) {
        std::cout << "move used\n"; // marker for EXPECT
        other.value = 0;
    }
    // Delete copy to emphasize move usage
    Counter(const Counter&) = delete;
    Counter& operator=(const Counter&) = delete;
    Counter& operator=(Counter&& other) noexcept {
        value = other.value;
        other.value = 0;
        return *this;
    }
};

int main(){
    Counter a{42};
    Counter b = std::move(a); // triggers noexcept move
    std::cout << "final b=" << b.value << "\n";
    return 0;
}

The program prints a marker that shows the noexcept move was used.

std::expected<T, E>

std::expected<T, E> represents either a value of type T or an error of type E. It is a typed, non‑throwing alternative for functions that can fail.

Compare to std::optional<T>: optional conveys only presence or absence. expected also conveys an error description.

Contrast to exceptions: expected returns an object that the caller must inspect. Control flow stays explicit in the source.

#include <iostream>
#include <string_view>
#include <string>
#include <expected>

std::expected<int, std::string> parse_int(std::string_view sv) {
    try {
        size_t pos = 0;
        int value = std::stoi(std::string(sv), &pos);
        if (pos != sv.size())
            return std::unexpected<std::string>("trailing characters");
        std::cout << "parsed\n"; // marker for EXPECT
        return value;
    } catch (const std::invalid_argument&) {
        return std::unexpected<std::string>("invalid integer");
    } catch (const std::out_of_range&) {
        return std::unexpected<std::string>("out of range");
    }
}

int main(){
    auto r1 = parse_int("123");
    if (r1) {
        std::cout << "value=" << *r1 << "\n";
    } else {
        std::cout << "error=" << r1.error() << "\n";
    }
    auto r2 = parse_int("abc");
    if (r2) {
        std::cout << "value=" << *r2 << "\n";
    } else {
        std::cout << "error=" << r2.error() << "\n";
    }
    return 0;
}

The demo parses an integer from a string view. On success it prints the value. On error it prints the error string.

std::expected composes. A function that calls another fallible function can pass the error upward with a single expression. .and_then chains success and .or_else handles failure. This keeps the error path linear and avoids the nested if/else pyramids that error-code APIs produce. The type carries both the value and the error, so the compiler tracks whether the result has been checked.

The monadic interface is what makes expected scale beyond two calls. Without it, a function that calls three fallible operations in a row needs three nested if statements, each checking and forwarding. With .and_then, the same logic is a single chained expression that reads left to right. The error short-circuits: the first failure skips the rest of the chain and returns the error to the caller. This is the same shape as Rust’s ? operator, expressed as method calls.

Contracts (C++26)

C++26 adds four contract attributes that can be placed on statements or function signatures.

  • [[assert]]: a debugging check that aborts if the condition is false.
  • [[assume]]: a promise to the optimizer that the condition is always true.
  • [[expects]]: a precondition that must hold when the function is entered.
  • [[ensures]]: a postcondition that must hold when the function returns.

The syntax is inline and looks like a standard attribute.

// Illustrative stub: does not compile on all compilers
[[expects: a > 0]]
[[ensures:  result >= a ]]
int factorial(int a) {
    [[assert: a >= 0]];
    int result = 1;
    for (int i = 2; i <= a; ++i) result *= i;
    return result;
}

Compilers differ in support. GCC 16 implements contracts behind a flag. Clang does not yet implement the feature. The snippet therefore serves only as an illustration.

Contracts encode the same invariants that the earlier chapters express with RAII and lifetimes. When the compiler cannot prove an invariant, a contract can document it and trigger a runtime check in debug builds.

Contracts and assertions serve different audiences. An assertion checks an internal invariant that the function itself controls. A precondition checks a contract between the caller and the callee: the caller guarantees the condition, and the callee relies on it. A postcondition runs the contract in reverse: the callee guarantees the condition, and the caller relies on it. The four attributes let you state who is responsible for what, which a plain assert cannot express.

The violation handler decides what happens when a contract breaks. The implementation can ignore the violation, log it, abort, or call a user-defined handler. The default is observe, which logs and continues. The enforce mode aborts. This flexibility lets a project ship with contracts in debug builds and strip them in release, or keep them on in production with a handler that reports to a monitoring service.

Historical perspective on error handling

Early C programs used integer return codes and errno as the sole mechanism for reporting failure. The programmer had to check each call manually, and forgetting a check produced silent corruption. C++98 introduced exceptions as a language feature that separates error propagation from the normal control flow. The first implementations used a table‑based “zero‑cost” model that stored unwind information in the object file. Compilers later added the ability to mark functions noexcept. This ability let the optimizer elide unwind handling for guaranteed‑non‑throwing code. C++23 added std::expected as a standard library type that returns either a value or an error object. Failure became an explicit part of the type system. C++26 formalizes contracts, completing the error‑handling toolbox.

Choosing a strategy

SituationRecommended tool
Failure cannot be recovered locally.Throw an exception.
Failure is routine and the caller must handle it inline.Return std::expected.
Absence carries no extra information.Return std::optional.
Invariant must never be violated.Use a contract attribute.
Function is guaranteed not to throw.Mark noexcept.

The table is a decision tree, not a menu. Read the failure mode first, then pick the tool. A function that fails because the input is bad returns expected. A function that fails because the system ran out of memory throws. A function that cannot fail marks itself noexcept. A function that documents a contract uses an attribute. The four tools cover four failure modes, and the modes do not overlap.

Use the tool that matches the semantic intent and avoid mixing strategies without a clear reason. A library that mixes expected and exceptions forces the caller to hold two mental models. Choose a single default per module and deviate only when the failure profile demands it. For example, a parser returns expected because every call can fail, while a memory allocator throws because out‑of‑memory is rare and unrecoverable.

Impossible cases

When the program reaches a state that the logic says cannot occur, the code can call one of the following.

  • std::abort(): terminates the process immediately.
  • std::terminate(): ends the program after invoking the terminate handler.
  • std::unreachable(): tells the compiler that this point is unreachable. Reaching it results in undefined behavior.

These calls are reserved for truly impossible branches such as a default case in a switch that covers every enum value.

std::unreachable is the strongest of the three. It tells the optimizer that the code path is dead, so the compiler can assume it never runs and use that fact to improve the surrounding code. If the assumption is wrong, the behaviour is undefined. Use it only after a construct that exhausts every case, such as a switch over a std::variant where the visitor covers every alternative. std::terminate is the safe fallback when the program is in a state you cannot reason about: it stops without running destructors, which avoids making things worse.

Try this

Write a function

std::expected<double, std::string> divide(double a, double b);

The function returns an error string when b == 0. Then write a caller that handles both the value and the error.

Text and formatting

In this chapter we replace the old C‑style formatting functions with the modern, type‑safe facilities that arrive in C++20 and C++23. The progression mirrors the lessons from the lifetime chapters: first we look at why printf is hazardous, then we introduce owning strings, non‑owning views, and finally the compile‑time checked formatting API.

The C printf problem

printf takes a format string that describes the types of the arguments that follow. The compiler cannot verify that the format string matches the argument list because the string is an ordinary char const *. If the programmer writes a mismatched conversion specifier or omits an argument, the program exhibits undefined behavior at run time.

// The format string expects an `int` and a `double`, but only one argument is supplied.
printf("%d %f\n", 42)

The above code compiles, but at execution the function reads a non‑existent double from the stack. The result is nondeterministic and leads to memory corruption.

std::string and std::string_view

std::string owns the character storage. It allocates memory, frees it when the object is destroyed, and therefore always remains valid for the lifetime of the owning object.

std::string_view does not own any characters. It merely holds a pointer and a length. The view is valid only while the underlying characters remain alive. Because a view never copies, it is ideal for function parameters: the caller can pass a std::string, a string literal, or a substring without allocating a new buffer.

void greet(std::string_view name) {
    std::println("Hello, {}!", name);
}

greet("Ada")                     // literal - no allocation
std::string full = "Grace Hopper"
greet(full)                       // temporary view of whole string
greet(full.substr(0,5))           // view of a prefix

The function greet never copies the characters it simply reads them through the view.

std::format (C++20)

std::format replaces printf with a type‑safe, variadic formatting function. The format string contains {} placeholders. The compiler parses the literal at compile time, matches each placeholder with the corresponding argument, and rejects mismatches with a diagnostic.

#include <format>
#include <print>
#include <string>

int main() {
    std::string name = "Alice";
    int age = 30;
    double score = 95.5;
    std::println("{}", std::format("Name: {}, Age: {}, Score: {:.2f}", name, age, score));
    return 0;
}

Running the program prints a line that contains Name: Alice. The EXPECT test in the build system checks for that substring.

std::print and std::println (C++23)

std::print writes to stdout without appending a newline. std::println does the same but adds a newline automatically. The book uses std::println for line-oriented output. It uses std::print only when the output is a sequence of values on one line, written inside a loop, where a newline after each value is wrong. In that case the loop body calls std::print("{} ", value) and a single std::println() closes the line after the loop.

#include <print>

int main() {
    std::println("{:>10} | {:>8}", "Item", "Value");
    std::println("{:>10} | {:>8.2f}", "A", 1.23);
    std::println("{:>10} | {:>8.2f}", "B", 4.56);
    std::println("{:>10} | {:>8.2f}", "C", 7.89);
    return 0;
}

The example prints a small table. The test harness looks for the token A in the output.

FacilityHeaderType safetyFormat checkingReturnsUse when
std::cout<iostream>per insertionnonestream referencestream features, custom operator<<
printf<cstdio>nonenoneint countC interop, legacy code
std::print / std::println<print>typed argumentscompile timevoidnew code, checked formatting

Format specifiers

The part of the format string after a colon controls alignment, width, fill character, precision, and numeric base.

  • Width and alignment: {:<10} left‑aligns inside a field of ten characters, {:>10} right‑aligns, {:^10} centers.
  • Fill character: {:*^8} pads with * while centering.
  • Precision: {: .2f} prints a floating‑point value with two digits after the decimal point.
  • Base: {:#x} prints an integer in hexadecimal with a 0x prefix, {:08b} prints binary padded to eight digits.
std::println("{:>10}", "right")          // "    right"
std::println("{:08b}", 5)                // "00000101"
std::println("{:#x}", 255)              // "0xff"
std::println("{:.2f}", 3.14159)          // "3.14"

If a specifier is unknown, the compiler issues an error because the format string is a compile‑time constant.

The specifier syntax mirrors Python’s format. The colon introduces a format‑spec, then optional fill, alignment, sign, alternate form, zero‑pad, width, precision, and type. The order matters. The compiler enforces it.

Performance of std::format is comparable to hand‑written printf for simple cases. The library avoids temporary allocations for short strings by using a small‑buffer optimisation.

Custom types can be formatted by specialising std::formatter. The specialization returns a format_to function that writes the representation into the provided output iterator.

struct Point { int x; int y; };

template<> struct std::formatter<Point> {
    constexpr auto parse(auto& ctx) { return ctx.begin(); }
    auto format(Point const& p, auto& ctx) const {
        return std::format_to(ctx.out(), "({},{})", p.x, p.y);
    }
};

std::println("{}", Point{3,4}); // prints "(3,4)"

The same custom formatter works for both std::println and std::format because they share the formatter protocol.

Advanced formatting features

Beyond the basic width and precision specifiers, std::format supports a rich set of options that let developers tailor the textual representation of values.

  • Sign handling: {: +} forces a leading plus sign for positive numbers, while {: -} (the default) prints a minus sign only for negatives.
  • Alternate form: The # flag adds a prefix for certain types: 0x for hexadecimal, 0 for octal, and a trailing decimal point for floating‑point values.
  • Zero padding: {:08} pads the field with zeros instead of spaces. It is equivalent to {:0>8} but more concise.
  • Grouping: {:L} formats numbers according to the locale’s thousands separator. This works together with a std::locale overload.
  • Date and time: When the <chrono> library provides a std::chrono::year_month_day or std::chrono::hh_mm_ss object, the formatter can emit ISO‑8601 strings using the {:T} or {:D} specifiers.
std::println("{:+08}", 42)        // "+0000042"
std::println("{:#x}", 255)       // "0xff"
std::println("{:L}", 1234567)    // "1,234,567" in en_US locale

using namespace std::chrono
auto now = floor<seconds>(system_clock::now())
std::println("{:T}", now)        // "2026-08-21T14:35:00"

These specifiers are composable a format string can combine alignment, fill, width, sign, and type in a single placeholder. The compiler checks that the combination is valid for the argument type, rejecting illegal mixes such as a sign flag on a string.

Custom types can also honour these flags by inspecting the format_context and the parsed format‑spec. A formatter can honour fill and align by delegating to std::format_to with a constructed format string, or it can implement the formatting logic directly for maximum performance.

struct Money {
    int cents;
};

template<> struct std::formatter<Money> {
    char presentation = 'f'; // f = dollars.cents, e = euros
    constexpr auto parse(auto& ctx) {
        auto it = ctx.begin();
        if (it != ctx.end() && (*it == 'e' || *it == 'f')) presentation = *it++;
        return it;
    }
    auto format(Myney const& m, auto& ctx) const {
        if (presentation == 'e')
            return std::format_to(ctx.out(), "€{:.2f}", m.cents / 100.0);
        else
            return std::format_to(ctx.out(), "${:.2f}", m.cents / 100.0);
    }
};
std::println("{}", Money{1234}) // prints "$12.34"

The example demonstrates how a formatter can respect a presentation specifier while still supporting the generic alignment and width options supplied by the surrounding format string.

Placeholders can refer to arguments by position.

std::println("{0} + {0} = {1}", 2, 4) // prints "2 + 2 = 4"

Named arguments are not part of the core standard, but a user can achieve them by providing a custom std::formatter specialization or by using std::make_format_args together with a format string that references names via a library such as {fmt}. The book mentions the technique briefly because it is useful in larger projects.

// Using a custom formatter (illustrative - not compiled here)
struct Point { int x int y }
template<> struct std::formatter<Point> : std::formatter<std::string> {
    auto format(Point const& p, auto& ctx) const {
        return std::formatter<std::string>::format(
            std::format("({},{})", p.x, p.y), ctx)
    }
}

Compile‑time checking is the point

std::format only accepts a compile‑time constant format string for full checking. If a program needs a run‑time format, the library provides std::runtime_format, which disables compile‑time verification. This design forces the programmer to choose safety whenever possible.

Contrast the two approaches:

  • printf("%d %f\n", 42): compiles, can crash at run time.
  • std::format("%d %f\n", 42): fails to compile because % is not a valid placeholder.
  • std::format(std::runtime_format(fmt), 42): compiles, but the format string is unchecked.

The book’s theme is to let the compiler catch what it can, and std::format embodies that principle.

Performance and constexpr formatting

std::format is designed to be fast. The implementation parses the format string at compile time when the literal is a constant expression, eliminating runtime parsing overhead. This makes it comparable to hand‑written printf for simple cases while providing safety.

When the format string cannot be known at compile time, the library falls back to a runtime parser. The cost is modest: a single pass over the format string plus the usual formatting work. For tight loops where every nanosecond matters, developers can still write a custom printf‑style loop, but the safety trade‑off must be justified.

C++23 extends std::format with constexpr support. A constexpr function can call std::format to produce a compile‑time constant string that can be used as a non‑type template parameter or in a static_assert.

constexpr std::string_view make_label(int id) {
    return std::format("Item-{:03}", id)
}
static_assert(make_label(7) == "Item-007")

The example illustrates how formatting can participate in compile‑time computation, enabling expressive metaprogramming without sacrificing safety.

Locale‑aware formatting is optional. By default std::format uses the “C” locale, which formats numbers with a period as the decimal separator. Passing a std::locale object to the overload selects the suitable digit grouping and decimal marks for the target locale.

std::locale german("de_DE")
std::println(std::format(german, "{:L}", 1234567.89)) // prints "1.234.567,89"

Custom formatters, shown earlier, work uniformly with locale‑aware overloads because the formatter receives the locale via its format method.

Error handling in formatting

When a format string is ill‑formed, std::format throws a std::format_error. This exception type is a subclass of std::runtime_error and carries a message that identifies the problem, for example, “argument index out of range” or “invalid format specifier”.

try {
    std::println(std::format("{0} {2}", 1, 2)) // index 2 does not exist
} catch (const std::format_error& e) {
    std::cerr << "Formatting failed: " << e.what() << '\n'
}

The library also provides the non‑throwing overload std::format_to_n which writes into a pre‑allocated buffer and returns the number of characters written. This is useful in low‑latency or embedded contexts where exceptions are disabled.

char buf[32]
auto result = std::format_to_n(buf, sizeof(buf), "{:04x}", 0x1A3)
std::println("{} characters written", result.out - buf)

result.out points just past the last character written, allowing the caller to construct a std::string_view without an extra copy.

The library also defines std::vformat and std::vprint which take a std::format_args object generated by std::make_format_args. These functions enable runtime‑determined argument lists while still performing compile‑time checks on the format string itself.

auto args = std::make_format_args(42, 3.14)
std::println(std::vformat("int={}, double={}", args))

If the format string itself is not a constant expression, the parser cannot validate the placeholders against the arguments at compile time. In that case, the same run‑time checks apply, and mismatches still result in a std::format_error.

The distinction between compile‑time guarantees and run‑time safety mirrors the earlier discussion of printf. By default the library favours compile‑time safety developers can opt into run‑time flexibility when the application requires it.


Try this

Write a program that prints a three‑row table. Each column must have a fixed width. The second column contains floating‑point numbers printed with two decimal places and right‑aligned.

// Insert your own code here - use std::println and format specifiers.

When you run the program, the output must look like a tidy table with aligned columns.


NOTE: All examples in this chapter are compiled and tested with the Clang 22 toolchain using -std=c++26. The book_example macro registers each source file as a test target, and the EXPECT strings verify that the output contains the expected fragments.

Containers

Containers provide storage for collections of objects. The Standard Library supplies a family of containers that differ in allocation strategy, ordering guarantees, and performance characteristics. Choose a container that matches the required usage pattern. The default choice is std::vector because it stores elements contiguously, which gives good cache locality and enables constant‑time random access.

std::vector: the default container

std::vector<T> owns a dynamically allocated array of T. Elements are stored next to each other in memory. The implementation allocates a capacity that can be greater than the current size. When the size grows past the capacity, the vector allocates a new, greater block, copies or moves existing elements, and frees the old block. This reallocation invalidates all pointers, references, and iterators that refer to the previous storage.

Because reallocation can be expensive, two techniques reduce its impact:

  • Reserve: call reserve(10) before inserting ten or greater number of elements. The vector allocates space for at least n elements up front, so later push_back or emplace_back calls cannot trigger a reallocation.
  • Emplace: emplace_back(args…) constructs a new element directly in the storage. The arguments are forwarded to the element’s constructor, avoiding an extra copy or move.

Both techniques tie back to earlier chapters. Chapter 07 described the lifetime‑safety analysis that flags dangling pointers after a reallocation. Reserving ahead of time eliminates that risk for the most common case. Chapter 08 explained why passing arguments by value or reference matters. emplace_back constructs in place, satisfying the principle of “construct where you use”.

The vector grows geometrically, typically doubling capacity. This yields amortised constant‑time push_back. Use reserve when the final size is known to avoid intermediate reallocations and iterator invalidation.

Erasing an element from the middle of a vector shifts every later element down by one, so removal is O(n) in the number of elements after the erased position. erase and the C++20 erase/erase_if free functions return the new logical end. When you only need to drop the last few elements, pop_back or resize is cheaper.

std::vector<bool> is a partial specialization that packs bits, so operator[] returns a proxy object rather than a bool&. This breaks generic code that expects a real reference, and it is the one standard container that is not a true container. Prefer std::vector<char> or std::bitset when you need a sequence of individual bit values with normal reference semantics.

#include <vector>
#include <print>

int main() {
    std::vector<int> v;
    v.reserve(10);
    v.emplace_back(1);
    v.emplace_back(2);
    v.emplace_back(3);
    std::println("vector size: {}", v.size());
    return 0;
}

std::array: fixed size on the stack

std::array<T, N> stores exactly N objects of type T. The size is known at compile time, and the array lives in the surrounding object’s storage, which is the stack for local variables. No dynamic allocation occurs.

A classic C array (T a[N]) decays to a pointer when passed to a function, losing its size information. std::array retains its size via the member function size(). It also provides standard container interfaces (begin(), end(), operator[]), so generic algorithms work uniformly.

#include <array>
#include <print>

int main() {
    std::array<int,5> a{{1,2,3,4,5}};
    std::println("array size: {}", a.size());
    std::println("first element: {}", a[0]);
    return 0;
}

The example fills an std::array<int,5> with the values 1 through 5, prints the size, and prints the first element.

std::array is an aggregate, so it can be copy‑assigned and returned by value with no hidden allocation, and its size is part of the type. Prefer it over std::vector when N is small and fixed, because the data sits in the parent object with no pointer chase. Reach for std::vector once the size is dynamic or larger than a few dozen elements.

Sequence containers with different insertion properties

ContainerTypical useInsertion costIterator invalidation
std::dequePush or pop at both ends without moving existing elementsAmortised constant time at either endInserting at the front or back does not invalidate existing iterators. Inserting in the middle can invalidate.
std::listFrequent insertion or removal in the middle of a long listConstant time anywhereNo iterator is invalidated by insertion or removal, except for the iterator that is removed.
std::forward_listSingly‑linked list, minimal memory overheadConstant time insertion after a known iteratorSame rules as std::list. No iterator invalidation except for erased elements.

All three store elements non‑contiguously. The lack of cache locality makes them slower for tight loops that iterate over a large number of elements. Use them only when the insertion pattern outweighs the cache penalty.

// Example of a deque that pushes at both ends.
std::deque<int> dq;
for (int i = 0; i < 5; ++i) dq.push_back(i);
for (int i = 5; i < 10; ++i) dq.push_front(i);

A std::deque stores elements in fixed‑size chunks rather than one block, which is why pushing at the front never moves the existing elements and never invalidates their iterators. The trade‑off is a double indirection on access and a larger per‑element overhead than vector.

Ordered associative containers: std::map and std::set

std::map<Key, Value> and std::set<Key> are implemented as red‑black trees. They keep keys in sorted order defined by operator< (or a custom comparator). Lookup, insertion, and removal take O(log n) time. Iterators remain valid across insertions and deletions, except when the element itself is erased.

Use these containers when ordered iteration, range queries, or stable iterator validity are required.

Use operator[] to find or insert, and at() when a missing key must raise std::out_of_range instead of creating a default. The ordering follows operator< on the key by default. A custom comparator changes both the order and the equality test. std::multimap and std::multiset permit duplicate keys when that is needed.

std::set stores keys only and is the right choice when you need a sorted, deduplicated collection rather than key‑value pairs. Both map and set gain a contains member in C++20 that tests membership without constructing an iterator.

std::map<std::string, int> word_counts;
word_counts["apple"] = 3;
word_counts["banana"] = 5;
for (const auto& [w, c] : word_counts) {
    std::println("{}: {}", w, c);
}

Unordered associative containers: std::unordered_map and std::unordered_set

std::unordered_map<Key, Value> and std::unordered_set<Key> are hash tables. Average‑case lookup, insertion, and removal are constant time. The containers do not preserve any ordering of keys. The key type must be hashable. The standard library provides std::hash for fundamental types and for std::string. Custom types need a specialization of std::hash and an equality operator.

Because hash tables store elements in buckets, iterator invalidation rules differ from ordered containers: inserting does not invalidate iterators, but rehashing (which can occur when the load factor exceeds a threshold) invalidates all iterators.

An unordered_map has a higher per‑operation constant cost than a vector search over tiny sets, and it allocates buckets, so prefer it only once the keyed lookup actually pays off. reserve and max_load_factor let you tune the bucket count before a bulk insert.

Iteration order over an unordered_map is not specified and can change between runs, so never depend on it. The contains member tests membership in average constant time.

std::unordered_map<int, std::string> id_to_name{{1, "Alice"}, {2, "Bob"}};
if (auto it = id_to_name.find(1); it != id_to_name.end()) {
    std::println("Found {}", it->second);
}

std::span: a non‑owning view over contiguous data

A std::span<T> is a lightweight object that refers to a contiguous sequence of T. It does not own the elements. It only stores a pointer and a length. Because it is non‑owning, it can be created from multiple sources: a std::vector<T>, a std::array<T, N>, or a raw C array.

Using std::span allows algorithms to accept any of those sources without copying. The container’s lifetime must outlive the span. Otherwise the span becomes dangling, which the lifetime‑safety analysis flags.

#include <vector>
#include <array>
#include <span>
#include <print>

void print_span(std::span<const int> s) {
    std::print("span elements:");
    for (int v : s) {
        std::print(" {}", v);
    }
    std::println("");
}

int main() {
    std::vector<int> v{1, 2, 3};
    std::array<int,3> a{{4,5,6}};
    print_span(v);
    print_span(a);
    return 0;
}

The print_span function takes a std::span<const int> and prints each element. The main function calls it with a std::vector<int> and an std::array<int,3>, demonstrating the uniform interface.

A std::span can carry a compile‑time extent (std::span<T, N>) or a dynamic one. The static form lets the compiler prove bounds in some algorithms. first, last, and subspan produce new spans over subranges without copying, which is how zero‑copy parsing pipelines stay allocation free.

A span exposes size, empty, and data, and it converts to a std::vector only through an explicit constructor, never by accident. This keeps ownership explicit at every call site.

Choosing the right container

When you start a new piece of code, follow this decision guide:

  1. Begin with std::vector. Its contiguous storage gives the best cache performance for most workloads.
  2. Switch to std::array if the number of elements is known at compile time and the total size fits within typical stack limits (e.g., ≤ 1 KB).
  3. Select an associative container when you need O(1) average‑case lookup by key. Choose std::map if you require ordered iteration or range queries. Choose std::unordered_map for constant‑time performance when ordering does not matter.
  4. Consider std::deque only when you need efficient insertion or removal at both ends and cannot accept the iterator invalidation of a vector during growth.
  5. Use std::list or std::forward_list only when you must insert or erase frequently in the middle of a large sequence and the cache penalty is tolerable.
  6. Use std::span to write generic algorithms that operate on any contiguous view without taking ownership.

Prefer contiguous, owning containers such as std::vector (default) or std::array when size is fixed. Use std::span for non‑owning views and std::pmr allocators to customise allocation without changing the container type. The same cache‑locality reasoning applies to std::string, while std::string_view follows the view pattern described earlier.

Try this

Write a program that:

  1. Declares a std::vector<int>.
  2. Calls reserve(10).
  3. Uses emplace_back to add the values 1, 2, 3.
  4. Creates a std::span<int> that references the vector.
  5. Computes and prints the sum of the elements in the span.

No solution is provided. The reader must fill in the code.

Algorithms are the loops

A hand-written loop hides intent. Chapter 04 showed that a loop can miss an off-by-one error and can expose lifetime bugs when a pointer is taken to an element that later moves. A named algorithm states exactly what happens: find the value, count the matches, transform each element. The reader understands the code without inspecting the body.

Algorithms operate on iterator pairs, making them generic over containers and element types. The same call works for std::vector, std::array, std::list, or any range providing iterators, and for int, std::string, or a user type with operator<. This uniformity removes boilerplate, reduces bugs, and lets one algorithm serve every container.

Every standard algorithm carries a documented complexity guarantee. std::find and std::count_if run in linear time because they can inspect each element. std::sort guarantees O(n log n) comparisons in the worst case. Knowing these bounds helps you choose the right tool on a performance-critical path. If you need only the smallest element, std::min_element is cheaper than a full sort because it stops after a single linear scan.

The iterator pair model

All standard containers expose begin() and end(). They define a half-open range [begin, end). The range includes the element pointed to by begin and excludes the element pointed to by end. This convention lets algorithms stop exactly at the last element without an extra check.

#include <vector>
#include <algorithm>
#include <print>

int main() {
    std::vector<int> v{1,4,7,10};
    int target = 7;
    auto it = std::find(v.begin(), v.end(), target);
    if (it != v.end())
        std::println("found {}", *it);
    else
        std::println("not found");
    return 0;
}

The program creates a std::vector<int>, calls std::find, and prints whether the target value was found.

Non-modifying sequence algorithms

The library provides many read-only algorithms. std::find returns an iterator to the first element equal to a value. std::count returns the number of elements equal to a value. The trio std::all_of, std::any_of, and std::none_of evaluates a predicate over a range. std::count_if counts elements that satisfy a predicate.

Beyond std::find, std::count, and the all_of family, the library offers std::mismatch to find the first differing position between two ranges, std::equal to test equality, and std::search to locate a subrange. These algorithms never alter the container, so you can call them on a const object.

std::find_if takes a predicate instead of a value, locating the first element that satisfies a condition. std::find_first_of finds the first element that matches any value from a second range. These variants cover the common cases where you search by property rather than by equality.

std::for_each applies a callable to each element. It is the algorithm-shaped alternative to a raw range-for loop, and it makes the intent visible at the call site. std::adjacent_find locates the first pair of neighbouring elements that satisfy a condition, which is useful for detecting duplicates or trends in a sequence.

Modifying algorithms

Algorithms that write to a destination include std::copy, std::transform, std::fill, and std::replace. They accept iterator pairs for the source and destination. std::transform applies a unary operation to each source element and writes the result to the destination range.

#include <vector>
#include <algorithm>
#include <print>

int main() {
    std::vector<int> v{1, 2, 3, 4, 5};
    std::vector<int> out(v.size());
    std::transform(v.begin(), v.end(), out.begin(), [](int x){ return x * x; });
    std::println("squared:");
    for (int n : out) std::print("{} ", n);
    std::println("");
    return 0;
}

The program prints the squared numbers. The destination must have room for every written element. std::back_inserter grows the container as needed.

std::remove and std::remove_if shift the kept elements to the front and return a new logical end. They do not erase anything. The erase-remove idiom combines the two: call erase with the iterator pair the algorithm returns to drop the unwanted tail in one statement. This avoids building a second container and reuses the original storage.

Beyond transform and remove, the library offers std::fill to assign a value to every element, std::replace to swap one value for another, std::rotate to cycle a subrange, and std::partition to group elements that satisfy a predicate before those that do not. Each returns the iterator or range you need to continue working without re-scanning.

std::unique removes consecutive duplicates, so it is effective only on a sorted range. Pair it with std::sort to drop all duplicates, then erase the trailing run as with remove.

std::sort rearranges elements into ascending order using operator<. std::stable_sort preserves the relative order of equal elements. After sorting, binary search algorithms become valid. std::binary_search reports whether a value exists in a sorted range. std::lower_bound returns the first position where a value can be inserted without breaking order. std::upper_bound returns the position after the last equal element.

The precondition matters. Running std::lower_bound on an unsorted range yields undefined results. The algorithm assumes monotonic ordering and will silently produce the wrong answer.

#include <vector>
#include <algorithm>
#include <print>

int main() {
    std::vector<int> v{5, 2, 9, 1, 5, 6};
    std::sort(v.begin(), v.end());
    std::println("sorted:");
    for (int n : v) std::print("{} ", n);
    std::println("");
    int key = 5;
    auto it = std::lower_bound(v.begin(), v.end(), key);
    if (it != v.end() && *it == key)
        std::println("lower_bound of {} is at index {}", key, std::distance(v.begin(), it));
    else
        std::println("key not found");
    return 0;
}

The program sorts a vector and uses std::lower_bound to locate the first occurrence of a value.

When you need only part of the order, do not pay for a full sort. std::partial_sort orders the first K elements and leaves the rest unspecified. std::nth_element places the Kth element in its sorted position and partitions the rest around it, all in linear time. These are the right tools for top-K and median queries.

Prefer std::stable_sort when equal elements carry order-dependent meaning, such as log entries that must stay chronological. The extra cost is small and the guarantee prevents subtle bugs when the sorted result feeds another pass.

Ranges (C++20)

C++20 introduced std::ranges. Ranges remove the need to pass iterator pairs. An algorithm can operate directly on a range object. The pipe syntax | composes adaptors that transform or filter the data before a terminal algorithm consumes it. Views are lazy: no intermediate container is allocated, and a pipeline can stop early.

#include <vector>
#include <ranges>
#include <print>

int main() {
    std::vector<int> v{1,2,3,4,5,6};
    auto pipeline = v
        | std::views::filter([](int x){ return x % 2 == 0; })
        | std::views::transform([](int x){ return x * x; });
    std::println("even squares:");
    for (int n : pipeline) std::print("{} ", n);
    std::println("");
    return 0;
}

The pipeline filters even numbers, squares them, and prints the result.

The adaptor set is larger than filter and transform. std::views::take keeps the first N elements, std::views::drop skips them, std::views::reverse inverts order, and std::views::split breaks a range on a delimiter. Because each view is lazy, v | std::views::filter(f) | std::views::take(3) examines elements only until three match. A ranges algorithm returns a view or subrange, so the result can feed another pipeline directly.

std::views::iota generates a numeric sequence without storing it, so std::views::iota(0, n) replaces a hand-written counter loop. Combined with filter and transform, it builds lazy numeric pipelines that allocate nothing. Ranges algorithms also accept projections on the terminal call, so std::ranges::sort(v, {}, &Point::x) sorts by the x member directly. std::ranges::to materialises a view into a concrete container when you finally need ownership, for example auto v = range | std::views::filter(f) | std::ranges::to<std::vector>(). This keeps the pipeline lazy until the boundary where storage is required.

Projection and comparator

std::ranges algorithms accept a projection argument. A projection extracts a member or computes a value before the algorithm compares or orders elements. This removes the need for an explicit comparator lambda. For example, sorting a vector of Point structs by the y coordinate needs no lambda:

struct Point { int x; int y; };
std::vector<Point> pts = {{1,5},{2,3},{4,7}};
std::ranges::sort(pts, {}, &Point::y);

The projection &Point::y tells the algorithm to compare the y members directly. The same idea applies to std::ranges::unique or std::lower_bound when you compare on a particular attribute.

Reduction

For many problems you need to combine a sequence of values into a single result. std::accumulate takes a beginning iterator, an ending iterator, and an initial value, then applies a binary operation (addition by default) to combine each element with the running total. std::ranges::fold_left is the ranges equivalent.

#include <vector>
#include <numeric>
#include <print>

int main() {
    std::vector<int> v{1,2,3,4,5};
    int sum = std::accumulate(v.begin(), v.end(), 0);
    std::println("sum = {}", sum);
    return 0;
}

The example builds a vector of five integers and computes their sum.

std::reduce is the parallel-friendly sibling of std::accumulate. It permits reordering of the operations, which lets an execution policy split the work across cores, but it requires the operation to be associative and the initial value to be an identity. Use std::accumulate when order matters and std::reduce when you only need the combined value.

std::inner_product combines two ranges with two operations. It multiplies corresponding elements and adds the products, which computes a dot product in one call. This is the reduction form of a zip operation, and it shows how a single algorithm can express what is otherwise a nested loop. The two-operation form generalises to any pair of associative combiners, so it can compute weighted sums or concatenations across two sequences in a single pass.

Algorithmic design patterns

Common tasks compose standard algorithms. Use filter‑map‑reduce: std::views::filter, std::views::transform, then std::accumulate or std::ranges::fold_left. Apply the erase‑remove idiom to discard elements without extra storage. Recognising these patterns yields concise code that leverages library guarantees.

Parallel policies

Many algorithms accept an execution policy as the first argument. std::execution::par asks the implementation to run the work in parallel when it can. Support for parallel policies is optional in the standard and depends on the compiler and its runtime backend. On some toolchains std::execution::par is unavailable or falls back to sequential execution. Treat parallel overloads as a performance option, not a correctness feature. When you use them, remember that order-unstable algorithms can reorder equal elements, so tests must check value-level properties such as sums rather than exact sequence.

std::vector<int> v = {3, 1, 4, 1, 5};
std::sort(std::execution::par, v.begin(), v.end());

Common pitfalls

  • Binary search on an unsorted range is undefined behaviour. Sort first.
  • Invalidated iterators. std::sort can invalidate all iterators. Reacquire them after sorting.
  • Iterator category mismatch. std::sort requires random‑access iterators. std::list iterators fail to compile.
  • Destination too small. std::copy and std::transform write exactly as many elements as the source supplies. Use std::back_inserter or size the destination first.
  • Sorting a node‑based container. std::list lacks random‑access iterators. Use its member list::sort instead.
  • Projection side effects. A projection must be pure. State‑modifying projections break algorithm invariants.

Try this

Use std::ranges to keep only the even elements of a std::vector<int>, square them, sort the result, and print each number on a single line.

// Write your solution here.

No solution is provided. The reader must fill in the code.

Ranges and views

What a view is

A view is a lightweight, non-owning handle over a sequence. It borrows the elements and never allocates memory. The view does not manage the lifetime of its elements. The same principle underlies std::string_view and std::span that were introduced earlier. Both types expose a pointer and a length, and they refuse to copy the data. A range is any object that provides begin() and end() that return iterators. A view is a range that does not own its elements. In other words, every view is a range, but not every range is a view.

Because a view never allocates, construction is a constant‑time pointer‑plus‑size operation. The compiler can inline it and no heap traffic occurs. The view inherits the lifetime constraints of its underlying storage, so a dangling std::span results if the source is destroyed, which the compiler warns about. Declaring a parameter as std::span<const T> also signals that the function reads elements without taking ownership, mirroring std::string_view for read‑only text. This combination of zero‑allocation construction and intent‑driven design reduces accidental copies and clarifies ownership boundaries across API surfaces.

A view has a precise definition in the standard. A type is a view if it is a range, is cheap to copy or move, and does not own the elements it presents. std::ranges even provides std::ranges::owning_view to wrap an owning container into the view model when an API demands a view but you must keep ownership locally. The inverse, std::ranges::ref_view, wraps a reference to a range you already own. Knowing which wrapper applies prevents both dangling and accidental copies.

std::views adaptors

In the <ranges> header, a namespace std::views provides a family of adaptors. Each adaptor returns a new view that lazily transforms the underlying range. The most useful adaptors are listed below.

  • filter: keeps only elements that satisfy a predicate.
  • transform: applies a function to each element.
  • take: stops after a given number of elements.
  • drop: skips a given number of elements.
  • reverse: iterates the range in reverse order.
  • split: splits a range on a delimiter.
  • iota: generates an infinite arithmetic progression.

These adaptors compose by piping (|). The pipe operator forwards the left-hand side as the source range to the right-hand adaptor, which returns a new view. Because each adaptor is itself a view, the resulting expression is a chain of lightweight objects that never allocate until the final consumption point.

The following example builds a pipeline that generates the integers from 1 to 10, keeps the even numbers, squares them, and takes the first three results. It then prints the numbers.

#include <ranges>
#include <iostream>
#include <vector>

int main() {
    // Generate numbers 1..10, keep evens, square them, take first three.
    auto rng = std::views::iota(1, 11)                 // 1..10 (exclusive upper bound)
               | std::views::filter([](int x){ return x % 2 == 0; })
               | std::views::transform([](int x){ return x * x; })
               | std::views::take(3);
    for (int v : rng) {
        std::cout << v << ' ';
    }
    std::cout << '\n';
}

It prints the three even squares 4 16 36. The EXPECT string in the build file verifies that these three numbers appear in the output.

Beyond the basic adaptors, the standard library provides combinators such as views::split for tokenising a string on a delimiter and views::reverse for reverse iteration without copying. The adaptor set also covers structure: views::chunk groups elements into fixed-size subranges, views::slide produces overlapping windows, and views::elements extracts the Nth member of each tuple‑like element (e.g., views::elements<0> extracts keys from a range of pairs). For associative containers, views::keys and views::values expose just the key or mapped type without copying. These utilities replace verbose loops and temporary containers with concise pipelines.

Laziness

Views are evaluated on demand. The pipeline does not create a temporary container after each adaptor. The take(3) adaptor stops the source after three elements have been produced. This property enables short-circuiting of expensive sources.

Consider a situation where the source is an unbounded range, such as std::views::iota(0). Without laziness, materialising the whole range requires infinite memory and never terminates. The lazy view evaluates each element only when the downstream consumer asks for it, and the take adaptor caps the evaluation.

The next example constructs an endless iota range, squares each value, and then takes only the first five results. Even though the source can produce an infinite number of elements, the program terminates after five squares because the view stops early.

#include <ranges>
#include <iostream>

int main() {
    // Infinite iota, square each, take first five elements.
    auto rng = std::views::iota(0)               // 0,1,2,... infinite
               | std::views::transform([](int x){ return x * x; })
               | std::views::take(5);
    for (int v : rng) {
        std::cout << v << ' ';
    }
    std::cout << '\n';
}

The output contains the first five squares 0 1 4 9 16. In contrast, a pre-ranges algorithm that first copies the iota range into a std::vector allocates billions of elements before the program must stop. The lazy view avoids that allocation entirely, saving both time and memory.

Laziness also improves cache behaviour. Each element is produced, transformed, and consumed in a single pass, keeping data in registers and avoiding a second memory pass. The same laziness applies to conditional stops: views::take_while keeps elements while a predicate holds and then halts, and views::drop_while discards until the predicate first fails. Because the adaptors are stateless, chaining drop_while with take_while on a sorted range extracts a contiguous band in one pass without building it.

Materialising a view: std::ranges::to

Sometimes code needs ownership of the elements produced by a view. The helper std::ranges::to materialises a view into a concrete container. The syntax is view | std::ranges::to<Container>(). The container type must be default-constructible and support push_back or equivalent insertion.

Materialisation is useful when an algorithm later requires random access, when the data must outlive the original source, or when an API expects an owning container. The operation copies each element exactly once and respects the allocator of the target container.

The target need not be a std::vector. std::ranges::to accepts any container meeting the insertion requirements, including std::list, std::deque, or a fixed-size std::array when the size is known. An optional allocator argument forwards to the container constructor, so materialisation can use a custom pool without changing the pipeline.

The following program creates the view iota(1,6), materialises it into a std::vector, and prints the size of the vector.

#include <ranges>
#include <iostream>
#include <vector>

int main() {
    auto rng = std::views::iota(1, 6); // 1,2,3,4,5
    auto vec = rng | std::ranges::to<std::vector<int>>();
    std::cout << "size = " << vec.size() << '\n';
}

The program prints size = 5, confirming that the view was copied into a container of the requested size.

Projections everywhere

Most range adaptors and algorithms accept a projection, a callable that extracts a member from each element. Projections let you avoid writing a lambda for a common operation. The syntax &T::member is a pointer-to-member that the algorithm treats as a projection.

Projections improve compile-time readability and often enable better inlining because the compiler sees a direct member access instead of a generic lambda capture. They also integrate with concepts that require a std::indirectly_readable predicate. This lets the standard library reason about the operation without executing user code.

The example below defines a simple Point struct with members x and y. A std::vector<Point> is sorted by the x coordinate using the projection &Point::x. The program then prints the sorted x values.

#include <algorithm>
#include <iostream>
#include <vector>

struct Point {
    int x;
    int y;
};

int main() {
    std::vector<Point> pts{{3,5},{1,2},{2,4}};
    // Sort by x using a projection
    std::ranges::sort(pts, {}, &Point::x);
    for (const auto& p : pts) {
        std::cout << p.x << ' ';
    }
    std::cout << '\n';
}

The output 1 2 3 demonstrates that the projection eliminated the need for a custom comparator.

Beyond sorting, any algorithm that accepts a projection can use the member‑pointer form. std::ranges::find(v, value, &T::key) searches a range of objects for a specific key without a bespoke lambda. The same member‑pointer projection works on associative ranges through views::values and views::keys. For example, m | std::views::values yields the integers from a std::map<std::string, int> directly. These patterns appear throughout the standard library and eliminate boilerplate code.

std::ranges algorithms versus pre-ranges algorithms

The range algorithms in std::ranges operate directly on any range, including views. They return iterators that refer to the original elements, so no copying occurs. The pipe syntax composes algorithms with adaptors. This makes the intent clear and the code compact.

The pre-ranges algorithmic style often required explicit iterator arguments, temporary containers, and separate calls to std::find or std::count. The range version can operate on a filtered view without materialising the intermediate sequence.

The next program filters a std::vector<int> for even numbers and then counts them using std::ranges::count_if. The count is printed.

#include <algorithm>
#include <iostream>
#include <vector>
#include <ranges>

int main() {
    std::vector<int> data{1,2,3,4,5,6};
    auto evens = data | std::views::filter([](int x){ return x % 2 == 0; });
    auto count = std::ranges::count_if(evens, [](int){ return true; }); // count elements in view
    std::cout << "evens = " << count << '\n';
}

The expected output evens = 3 shows that the algorithm works on the filtered view without materialising a new container.

Range algorithms also accept a projection argument, mirroring the adaptor behaviour. For example, std::ranges::sort(v, {}, &T::member) sorts a range of structures by a field without an explicit comparator lambda. This uniformity simplifies generic code that works with both plain values and aggregates.

Reduction joins the model as well. std::ranges::fold_left and std::ranges::fold_right accumulate a range with an initial value and a binary operation, returning the combined result rather than writing into a destination. They pair naturally with a filter and a transform, so the filter-map-reduce pattern from the algorithms chapter becomes one expression.

Custom views and std::range_adaptor_closure

Advanced users can create reusable pipelines by defining a range adaptor closure. A closure is a callable that returns a view when applied to a range. The closure can be used in a pipe expression just like the built-in adaptors.

Defining a closure typically involves a generic lambda that captures any configuration parameters and returns a composition of existing adaptors. Because the closure returns a closure object (std::range_adaptor_closure), the compiler treats it as another adaptor, preserving laziness.

The example below defines a simple adaptor twice that multiplies each element by two. The adaptor is expressed as a generic lambda that returns a transform view. The program creates an iota range, pipes it through twice, and prints the results.

#include <ranges>
#include <iostream>

// A reusable, pipeable adaptor that doubles each element.
struct Twice : std::ranges::range_adaptor_closure<Twice> {
    template <std::ranges::viewable_range R>
    auto operator()(R&& r) const {
        return std::views::transform(std::forward<R>(r), [](int x) { return x * 2; });
    }
};

inline constexpr Twice twice;

int main() {
    auto rng = std::views::iota(1, 5) | twice;  // 1..4 doubled -> 2 4 6 8
    for (int v : rng) {
        std::cout << v << ' ';
    }
    std::cout << '\n';
}

The output 2 4 6 8 confirms that the custom adaptor behaved like a built-in view. Readers can extend this pattern to more sophisticated pipelines, such as a prime_filter that composes filter with a deterministic primality test. Because the closure is a compile-time object, the optimizer can erase the intermediate layers entirely.

A view lives only as long as its source. Returning a view to a local container or temporary yields a dangling view, which the compiler warns about (see chapter 07). Because the pipeline composes at compile time, the compiler can collapse multiple adaptor layers into a single loop, achieving speed comparable to hand‑written loops while preserving readability.

Try this

Build a pipeline that takes std::views::iota(1, n), keeps only the multiples of 3, squares each kept value, takes the first 5 results, and prints them. The program must compile with the book’s standard settings and must run without allocating an intermediate container.

No solution is provided. The reader must write the code, register it as a book_example, and verify that the output contains five numbers.

Callables and type erasure

What is a callable

A callable denotes any entity that can be used after the function‑call operator (). The C++ standard groups several kinds under this umbrella: a function pointer to a free function, a function object (functor) that overloads operator(), a lambda (an unnamed function object generated from a capture list and body), std::function (a type‑erased wrapper that can hold any of the previous forms when the call signature matches), and a pointer to a member function or data member used with an object instance.

Generic algorithms (chapter 12) and range algorithms (chapter 13) accept a callable parameter expressed as a template type satisfying the invocable requirement, which allows the same algorithm to work with a raw pointer, a capturing lambda, or a std::function. This decouples the algorithm from the concrete call target and is central to generic programming.

Lambdas

Lambdas provide a concise way to create a function object. The syntax starts with a capture list in square brackets, followed by an optional parameter list, an optional mutable specifier, an optional exception specification, and a body. The capture list determines which surrounding variables become members of the closure type.

  • [x] copies the variable x into the closure. The copy is a separate object. Later modifications of the original x do not affect the captured value.
  • [&x] stores a reference to x. The closure can read and write the original variable through the reference.
  • [=, this] captures all automatic variables that appear in the body by copy and also captures the current object *this by reference. This form is useful inside a non‑static member function when the lambda needs to read both data members and local variables.
  • [] declares a capture‑less lambda. Such a lambda has no state and can be converted to a function pointer if it does not use this.

A lambda can be declared mutable to allow modification of its captured copies. Without mutable the operator() of the closure is const, which forbids changing any captured by copy. Since C++20 a lambda can introduce explicit template parameters using the abbreviated syntax []<typename T>(T x) { return x }. The compiler treats this as a generic lambda that can be instantiated with any type T that satisfies the body.

The following example demonstrates a lambda used as a predicate for std::ranges::count_if. The program builds a vector of integers, counts the elements that are even, and prints the result. The lambda captures nothing and therefore has the type bool(int) const.

#include <vector>
#include <iostream>
#include <algorithm>
#include <ranges>

int main(){
    std::vector<int> v = {1,2,3,4,5,6};
    auto is_even = [](int x){ return x % 2 == 0; };
    auto count = std::ranges::count_if(v, is_even);
    std::cout << "count = " << count << std::endl;
    return 0;
}

The test harness checks that the program prints the line count = 3. The lambda illustrates how a small, capture‑less callable integrates directly with a range algorithm.

Capturing state and performance

When a lambda captures by value, the closure holds its own copy, which makes the lambda safe to copy and move. Capturing by reference stores a reference member, so copies still refer to the original variable, which matters when the lambda outlives that variable. A capture‑less lambda can convert to a function pointer, removing indirection. If a lambda captures state and the closure type is known, the compiler can inline the call and eliminate the indirect call. The only overhead is copying captured values at construction.

std::function and type erasure

std::function<R(Args…)> is a class template that abstracts away the concrete callable type. Internally it stores a pointer to a type‑erased function object and a pointer to a virtual call dispatcher. When a callable is assigned to a std::function, the wrapper can allocate dynamic memory if the object does not fit into the Small‑Object Optimization buffer (typically 2 to 3 pointers). The allocation adds heap traffic and a level of indirection at each invocation.

The trade‑off is flexibility: std::function can hold any matching callable, which enables runtime polymorphism such as heterogeneous containers and user‑selected callables. In contrast, a template parameter preserves the concrete type, which allows the compiler to inline calls, eliminate the virtual dispatcher, and avoid heap allocation. This zero‑overhead approach is recommended for performance‑critical code.

The example below builds a std::vector<std::function<int(int)>>. It stores three different callables: a free function that multiplies its argument by three, a capturing lambda that adds five, and a std::bind expression that multiplies by five. The program iterates the vector, calls each element with the argument 5, and prints the result.

#include <vector>
#include <functional>
#include <iostream>
#include <utility>

int triple(int x) { return x * 3; }

int main(){
    std::vector<std::function<int(int)>> ops;
    ops.emplace_back(triple); // free function
    ops.emplace_back([](int x){ return x + 5; }); // lambda adds five
    using namespace std::placeholders;
    ops.emplace_back(std::bind([](int a, int b){ return a * b; }, _1, 5)); // multiply by 5

    for(const auto &op : ops){
        std::cout << op(5) << std::endl;
    }
    return 0;
}

The test expects the three lines 15, 10, and 25 in that order. The example shows how heterogeneous callables can coexist in a single container.

Allocation and when to prefer std::function

If the stored callable fits into the small‑object buffer, std::function does not allocate. In the example the lambda and the bound function are small enough, so no heap allocation occurs. If a callable captures a large std::vector by value, the wrapper will allocate to hold the captured data.

  • When the callable type is not known at compile time, such as when a plugin supplies a callback.
  • When the callable must be stored in a homogeneous container that outlives the point of creation.
  • When the API is a boundary that other languages or runtimes will call.

When to erase, when to template

Choosing type erasure or a template depends on when the callable is known.

  • If the algorithm is a library component that the user instantiates, the library must expose a template parameter (or a generic auto parameter) for the callable. This yields a bespoke instantiation for each caller, which allows the compiler to inline the call and generate optimal code.
  • If the algorithm is part of a runtime system, such as a GUI framework that stores user‑supplied callbacks, a networking library that registers event handlers, or a scripting engine that invokes user code, type erasure via std::function or std::move_only_function is appropriate.

A practical decision rule:

  1. Ask whether the callable can be known at compile time.
  2. If yes, write a template parameter. The compiler will emit a separate version for each distinct callable type.
  3. If no, use std::function (or the move‑only variant) to hide the type.

The rule helps avoid accidental performance loss in tight loops while still providing the flexibility needed at program boundaries.

std::invoke and std::invoke_r

std::invoke is a utility that uniformly calls several kinds of callables. It accepts a callable and a set of arguments, then dispatches to the appropriate call expression. The overload handles:

  • Regular function objects and function pointers.
  • Pointers to member functions, where the first argument is the object (or a reference, pointer, or smart pointer) on which to invoke the member.
  • Pointers to data members, where the result is the member value.

std::invoke_r<R>(f, args…) adds an explicit return‑type conversion, which forces the result to be converted to R before returning.

The following program defines a free function, a struct with a member function and a data member, and then calls each through std::invoke. The output demonstrates how the same helper covers all three cases.

#include <iostream>
#include <functional>
#include <string>

struct Obj {
    int value = 42;
    int get() const { return value; }
    int data = 99;
};

int free_func(int x) { return x; }

int main(){
    // free function via std::invoke
    std::cout << "free: " << std::invoke(free_func, 42) << std::endl;

    Obj o{ .value = 7, .data = 99 };
    // pointer to member function
    auto mem_fn = &Obj::get;
    std::cout << "member: " << std::invoke(mem_fn, o) << std::endl;

    // pointer to data member
    auto mem_data = &Obj::data;
    std::cout << "data: " << std::invoke(mem_data, o) << std::endl;
    return 0;
}

The test harness checks for the three lines free: 42, member: 14, and data: 99. Using std::invoke simplifies generic code because the same syntax works for all callable categories.

Function objects / functors

A function object, commonly called a functor, is a class or struct that defines operator(). The type is a regular class type, so it can have explicit constructors, data members, and custom copy or move behaviour. Functors are useful when the callable carries significant state that must be part of its type, for example when the state influences overload resolution or when the type must be named for ADL.

The example below defines a comparator that orders integers by their absolute distance from a reference value. The comparator stores the reference as a data member and implements a constexpr call operator. The program creates a vector, sorts it with std::ranges::sort using the functor, and prints the sorted sequence.

#include <vector>
#include <algorithm>
#include <iostream>
#include <ranges>
#include <cmath>

struct DistanceComparator {
    int ref;
    constexpr bool operator()(int a, int b) const {
        return std::abs(a - ref) < std::abs(b - ref);
    }
};

int main(){
    std::vector<int> v = {5,1,3,2,4};
    std::ranges::sort(v, DistanceComparator{0}); // sort by distance from 0 (i.e., absolute value)
    for(int x: v) std::cout << x << ' ';
    std::cout << std::endl;
    return 0;
}

The test checks that the program prints 1 2 3 4 5. The functor demonstrates how a stable type can be passed to algorithms without incurring any type‑erasure cost.

When a functor is preferable

  • When the callable must be copyable or movable in a predictable way.
  • When the functor participates in overload sets that depend on its type.
  • When the callable is part of a public API and the type name conveys intent.

std::move_only_function (C++23)

std::move_only_function<R(Args…)> is introduced in C++23 as the move‑only analogue of std::function. It provides type erasure while forbidding copying. This enables storage of callables that are not copyable, such as a lambda that captures a std::unique_ptr or a class that holds a mutex.

The move‑only wrapper eliminates the hidden copy‑construction step that std::function performs when a copy of the wrapper is made. When the program never copies the wrapper, the overhead is identical to std::function but with the added guarantee that the stored callable cannot be duplicated inadvertently.

Note: The book registers this feature as a gap on compilers that do not yet implement the header. The prose explains the semantics. The example is omitted to keep the build portable.

Try this

Build a std::vector<std::function<double(double)>> containing three operations: the identity function, a lambda that squares its argument, and a lambda that adds one. Invoke each stored callable with the value 3.0 and print the three results on separate lines.

No solution is provided. The reader must write the code, register it as a book_example, and verify that the output contains the three numbers 3, 9, and 4.

Numerics and multidimensional views

Type‑safe math constants: std::numbers

The header <numbers> supplies inline constexpr constants for float, double, and long double. The primary name, e.g. std::numbers::pi, denotes a double. The alias std::numbers::pi_v<T> yields the constant in type T. This removes the need for separate literals such as M_PI or user‑defined constexpr values.

The older macro M_PI originates from the C header <cmath>. It expands to a literal of type double and is not guaranteed to exist on all platforms. Because it is a macro, the constant cannot participate in overload resolution based on the target type. The std::numbers objects avoid those problems and can be used in all constexpr expressions.

The family includes more than π. std::numbers::e, std::numbers::sqrt2, std::numbers::ln2, and std::numbers::phi cover the constants most numerical code needs. The _v alias keeps the chosen precision: std::numbers::pi_v<float> is a single-precision approximation, while std::numbers::pi_v<long double> carries the widest precision the platform supports. Because the constants are constexpr, they can initialise static and consteval contexts without a runtime cost.

The header also provides reciprocals such as std::numbers::inv_pi and std::numbers::inv_sqrt2, which avoid a division at the call site when an inverse is what the math requires.

The example below prints the value of π, computes the area of a circle with radius 2.5, and formats the result with std::cout. The test harness checks that the output contains the word area.

#include <iostream>
#include <numbers>

int main(){
    double r = 2.5;
    double area = std::numbers::pi * r * r;
    std::cout << "area = " << area << std::endl;
    return 0;
}

Core floating‑point utilities: <cmath>

The header <cmath> implements the classic mathematical functions. All functions are overloaded for float, double, and long double. Since C++26 the overload set also accepts integral arguments, promoting them to the appropriate floating type.

  • std::sqrt(x) returns the square root of x.
  • std::pow(b, e) raises b to the power e.
  • std::hypot(x, y) computes √(x² + y²) while protecting against overflow and underflow. The expression std::sqrt(x*x + y*y) can overflow when the magnitude of x or y is large.
  • std::floor(x) returns the greatest integer not larger than x.
  • std::ceil(x) returns the smallest integer not smaller than x.

The set also contains safer primitives for interpolation and rounding. std::midpoint(a, b) returns the value halfway between a and b without the overflow that a + (b - a) / 2 can suffer. std::lerp(a, b, t) computes a + t * (b - a) with proper endpoint handling. std::fma(x, y, z) computes (x * y) + z as a single fused operation, avoiding an intermediate rounding step and improving speed and accuracy on supporting hardware.

The classification functions std::isfinite, std::isnan, and std::isinf test a value’s category without throwing, which provides a reliable way to detect a failed computation.

In addition, std::copysign(x, y) copies the sign of y onto the magnitude of x, which is the correct way to negate a zero or to preserve a sign across an operation. std::nextafter(x, y) steps to the next representable value toward y, exposing the discrete nature of floating-point for tolerance and unit-test work.

In the following program we compute a right‑angled triangle with legs 3 and 4. The call to std::hypot yields 5 without intermediate overflow. The test expects the line hypot = 5.

#include <iostream>
#include <cmath>
int main(){
    double a = 3.0, b = 4.0;
    double h = std::hypot(a, b);
    std::cout << "hypot = " << static_cast<int>(h) << std::endl;
    return 0;
}

Complex arithmetic: std::complex

The class template std::complex<T> stores a complex number whose real and imaginary parts are of type T. Member functions real() and imag() provide access to the components. The standard library overloads the arithmetic operators so that addition, subtraction, multiplication, and division operate component‑wise.

Two helper functions are useful for magnitude calculations. std::norm(z) returns the squared magnitude (real(z)² + imag(z)²). std::abs(z) returns the magnitude itself (√norm(z)). The factory function std::polar(r, θ) constructs a complex number from polar coordinates, where r is the radius and θ the angle in radians.

The library also overloads the transcendental functions for std::complex, so std::sin, std::exp, and std::log accept complex arguments and return complex results. std::conj(z) returns the conjugate, and std::proj(z) returns the projection onto the Riemann sphere, which matters for infinities. Only std::complex<float>, std::complex<double>, and std::complex<long double> are well formed. Instantiating std::complex<int> is ill formed, because integer complex arithmetic is not defined by the standard.

A user-defined literal makes complex literals readable. The std::literals::complex_literals inline namespace provides the i suffix, so auto z = 1.0 + 2.0i builds 1 + 2i without a constructor call. The literal works for float, double, and long double.

The program below constructs a complex number z = 3 + 4i, prints its real and imaginary parts, computes its magnitude with std::abs, and creates a second complex number via std::polar. The test looks for the word real in the output.

#include <iostream>
#include <complex>
#include <numbers>

int main(){
    std::complex<double> z(3.0, 4.0);
    std::cout << "real = " << z.real() << ", imag = " << z.imag() << std::endl;
    std::cout << "abs = " << std::abs(z) << std::endl;
    auto p = std::polar(2.0, std::numbers::pi/4);
    std::cout << "polar real = " << p.real() << std::endl;
    return 0;
}

Compile‑time rational numbers: std::ratio

std::ratio<N, D> encodes a rational number as two compile‑time integer template arguments. The type can be used in non‑type template parameters, which enables compile‑time arithmetic without additional run‑time cost.

The library provides metafunctions such as std::ratio_add<A, B> and std::ratio_multiply<A, B>. These compute a new std::ratio that represents the sum or product of the two operand ratios. The result is available as a nested type member.

std::ratio is the foundation of the <chrono> duration type. std::chrono::milliseconds is std::chrono::duration<int, std::ratio<1, 1000>>, so a ratio literal becomes a unit of time. The trait std::ratio_equal<A, B> and the ordering std::ratio_less<A, B> let templates reason about the relationship between two ratios at compile time, and std::ratio_divide<A, B> produces the quotient. All of these are constexpr, so the compiler evaluates them entirely during compilation.

For example, std::chrono::duration<long, std::ratio<1, 1000>> is exactly milliseconds: a duration that stores a count of long ticks where each tick is one thousandth of a second. Swapping the ratio to std::ratio<1, 1> yields seconds and to std::ratio<60, 1> yields minutes, all from the same template.

The example defines a base ratio representing one thousandth (std::ratio<1, 1000>), prints its numerator and denominator, and then uses std::ratio_add to add two such ratios, which produces 2/1000. The test harness searches for the exact string 1/1000.

#include <iostream>
#include <ratio>

int main(){
    using thousand = std::ratio<1,1000>;
    std::cout << thousand::num << '/' << thousand::den << std::endl;
    using sum = std::ratio_add<thousand, thousand>::type;
    // sum is 2/1000, but we only need to show the base ratio
    (void)sum{}; // suppress unused warning
    return 0;
}

Integer helpers: std::gcd and std::lcm

The functions std::gcd(a, b) and std::lcm(a, b) compute the greatest common divisor and the least common multiple of two integral values. Implementations use highly optimized code to execute the Euclidean algorithm efficiently. Writing these algorithms manually can produce slower code.

The binary Euclidean algorithm runs in logarithmic time, so even 64-bit arguments complete in a few dozen cycles.

Both functions are constexpr and return the common type of their arguments after the usual arithmetic conversions. std::gcd is the building block for reducing fractions and for the Euclidean distance checks used in number theory, while std::lcm sizes a buffer that must hold a whole number of repeats of two periodic signals. Reaching for the standard functions avoids the off-by-one and signedness bugs that hand-written versions attract.

The program below calculates gcd(48, 180) and lcm(48, 180). The expected output contains the two lines gcd = 12 and lcm = 720.

#include <iostream>
#include <numeric>

int main(){
    int a = 48, b = 180;
    std::cout << "gcd = " << std::gcd(a, b) << std::endl;
    std::cout << "lcm = " << std::lcm(a, b) << std::endl;
    return 0;
}

Multidimensional non‑owning views: std::mdspan

std::mdspan is a non‑owning view that maps a contiguous block of memory to a multi‑dimensional array. It extends std::span to support a compile‑time known number of dimensions. The template parameters are the element type and an extents description, which can be static, dynamic, or mixed.

Two layout policies exist. std::layout_right stores elements in row‑major order, which matches the layout of C‑style arrays and of std::vector. std::layout_left stores elements in column‑major order, which aligns with the conventions of Fortran and some linear‑algebra libraries. The layout determines the stride that the view applies when a multidimensional index is supplied.

The extents can be dynamic instead of static. std::extents<std::size_t, std::dynamic_extent, 4> fixes only the column count and takes the row count at construction, which suits matrices whose size is known at runtime. submdspan slices a view into a smaller view without copying, mirroring std::span::subspan from chapter 11. The accessor template argument controls how elements are read and written, so an mdspan can present a strided or even a non-contiguous layout while keeping the same multidimensional interface.

Because mdspan is the multidimensional sibling of std::span, the same rule from chapter 11 applies. The view borrows storage and must not outlive the container it observes, or the indexed access reads freed memory.

Like std::span, mdspan is constexpr friendly, so a small matrix held in static storage can be indexed inside a consteval context and checked by static_assert.

The example creates a std::vector<double> with twelve values. It then constructs a std::mdspan<double, std::extents<std::size_t, 3, 4>, std::layout_right> that views the vector as a 3 × 4 matrix. A nested loop prints each row on a separate line. The test verifies the three rows 0 1 2 3, 4 5 6 7, and 8 9 10 11.

#include <iostream>
#include <vector>
#include <mdspan>

int main(){
    std::vector<double> data(12);
    for (std::size_t i = 0; i < data.size(); ++i) data[i] = static_cast<double>(i);
    std::mdspan<double, std::extents<std::size_t, 3, 4>, std::layout_right> view(data.data());
    for (std::size_t r = 0; r < 3; ++r) {
        for (std::size_t c = 0; c < 4; ++c) {
            std::size_t idx = r * 4 + c; // row‑major stride
            std::cout << *(view.data_handle() + idx);
            if (c + 1 < 4) std::cout << ' ';
        }
        if (r + 1 < 3) std::cout << '\n';
    }
    return 0;
}

Linear algebra: std::linalg (C++26)

The header <linalg> adds a lightweight linear‑algebra library that operates directly on std::mdspan objects. Matrix‑matrix, matrix‑vector, and vector‑vector products are expressed as free functions that accept std::mdspan parameters. Because the functions work on views, no temporary storage is required.

Unlike the decades-old BLAS interface, std::linalg is generic over the element type and the layout, so the same call works on float, double, or a custom accumulator, and on row-major or column-major storage without rewriting.

All operations are constexpr where the underlying arithmetic permits compile‑time evaluation. This enables static analysis of small linear‑algebra expressions and permits their use in static_assert statements.

Support for <linalg> is not yet present in the default Clang toolchain used by the build pipeline. The book therefore presents the feature in prose only, with a callout that states the feature is a new addition in C++26 and will become available when compilers implement the header. Consequently, any example that includes <linalg> fails to compile until the compiler ships the header.

The functions follow a consistent naming pattern. linalg::matrix_product, linalg::dot, and linalg::vector_norm take read-only input views and an output view, returning the result through the last argument rather than by value. This output-parameter style keeps the operation allocation free and lets the caller choose the storage.

Try this

Create a std::vector<double> that contains twelve values of your choice. View the vector as a 3 × 4 std::mdspan. Print the element at row 1, column 2 using the call syntax mdspan(r, c). Verify that the printed value matches the element you stored.

Templates I: functions that match

Why templates

Templates let a single definition work for many types. The compiler creates a concrete version for each type that appears in a program. This mechanism is compile‑time polymorphism. Runtime polymorphism uses virtual functions, a v‑table, and an indirection. Templates eliminate both indirection and heap allocation. The standard library builds every container and algorithm from templates. std::vector<T> and std::ranges::sort illustrate this fact.

Read templates as term-rewriting rules with unification, in the sense a Prolog programmer knows. A template is a pattern. The compiler unifies the call’s argument types against the pattern’s parameters, binding T to a concrete type, then rewrites the call into a specialized instance. Overload resolution and specialization are rule precedence on top of that unification. This mental model explains why error messages name a failed unification rather than a line in your logic.

Templates let library authors write code that works for any iterator type, raw pointers, std::vector iterators, or user‑defined iterators. The compiler enforces required operations at each instantiation and reports errors attached to the generated code, providing early feedback. Because each specialization is generated at compile time, the optimizer can inline, eliminate dead code, and propagate constants, yielding zero‑overhead performance comparable to hand‑written code. This combination of type‑safe generic programming and compile‑time abstraction makes templates the preferred tool for reusable library components.

Function templates

A function template starts with the keyword template. The following example defines a generic max function.

template<typename T>
T max(T a, T b) {
    return a < b ? b : a;
}

The compiler deduces T from the arguments at the call site. When deduction fails, the caller can supply the type explicitly, for example max<int>(a, b). Deduction works for built‑in types, standard library types, and user‑defined types that provide the required operators.

Deduction can fail or surprise. max(1, 2.0) is ambiguous because T cannot be both int and double. The fix is an explicit max<double>(1, 2.0) or converting the arguments first. The same ambiguity appears whenever two arguments disagree in type, which is why generic helpers often take const T& and let the caller supply matching types.

Template argument deduction follows a set of deduction guides. When a parameter is a reference, the reference qualifiers are preserved. When a parameter is a forwarding reference (T&&), deduction yields an lvalue reference for lvalue arguments and an rvalue reference for rvalue arguments. This mechanism underlies perfect forwarding, a technique used throughout the standard library.

The example program calls max with two int values and prints the result.

#include <print>

template<typename T>
T max(T a, T b) {
    return a < b ? b : a;
}

int main() {
    int a = 3;
    int b = 9;
    std::println("max = {}", max(a, b));
}

Class templates

Class templates use the same syntax as function templates. The example defines a simple wrapper Box that stores a value of type T.

template<typename T>
struct Box {
    T value;
};

std::vector<T> (introduced in chapter 11) is the canonical class template. A class template can contain member functions, static data, and nested type definitions that also depend on the template parameters.

Class template argument deduction (CTAD) lets the compiler infer T from the constructor arguments, so Box(42) deduces Box<int> without writing the angle brackets. A deduction guide can teach the compiler a non-obvious mapping, for example deducing a Box<std::string> from a string literal. CTAD removes the boilerplate that earlier C++ required for every generic constructor call.

The example program creates a Box<int> and a Box<std::string> and prints each value.

#include <print>
#include <string>

template<typename T>
struct Box {
    T value;
};

int main() {
    Box<int> ibox{42};
    Box<std::string> sbox{"hello"};
    std::println("int box: {}", ibox.value);
    std::println("string box: {}", sbox.value);
}

A member variable inside a class template is instantiated once per distinct template argument. For instance, a class template can define int count to provide a separate count for each T. This pattern supports compile‑time registries with zero runtime allocation.

Abbreviated function templates (C++20)

C++20 introduced a shorthand for simple function templates. The declaration auto f(auto x) expands to template<typename T> auto f(T x). The following program demonstrates both forms.

// full form
template<typename T>
auto identity_full(T x) { return x; }

// abbreviated form
auto identity_abbrev(auto x) { return x; }

Both functions return their argument unchanged. The abbreviated form reduces boilerplate and improves readability when the function body does not depend on the template parameter name.

#include <print>

template<typename T>
auto identity_full(T x) { return x; }

auto identity_abbrev(auto x) { return x; }

int main() {
    std::println("full = {}", identity_full(7));
    std::println("abbrev = {}", identity_abbrev(7));
}

Abbreviated templates also integrate with generic lambdas. A lambda such as [](auto x){ return x } is internally a function template, allowing the same zero‑overhead semantics.

An auto parameter in a template is a forwarding reference when it is a template parameter: template<typename T> void f(T&& x) and void f(auto&& x) both bind T to the value category of the argument, so std::forward<T>(x) preserves whether the caller passed an lvalue or an rvalue. This is the mechanism behind std::make_unique and the perfect-forwarding wrapper from chapter 08. A plain T&& that is not a template parameter is an rvalue reference, not a forwarding reference, so the distinction lives in the template parameter list.

Templates also accept a variable number of type parameters with a parameter pack: template<typename... Ts> struct Tuple { }. The ellipsis ... both declares the pack and expands it. Fold expressions collapse a pack with a binary operator, so ((std::cout << args << ' ') ...) prints every argument without writing a loop or a recursive base case. Pack expansion is the compile-time counterpart of a variadic function, and it produces a fully inlined call sequence.

For example, a single function can sum a pack with (0 + ... + args), or print every argument with ((std::cout << args << ' ') ...). The compiler expands the pattern into std::cout << a << ' ' , std::cout << b << ' ' and so on, with no recursion and no runtime dispatch. This is the modern replacement for the recursive variadic helpers that older C++ required.

Dependent names and the typename/template keywords

When a name depends on a template parameter, the compiler cannot know whether it is a type or a value. The following snippet shows two common gotchas.

template<typename T>
void foo(T t) {
    // dependent type requires 'typename'
    typename T::type *ptr = nullptr;
    (void)ptr; // silence unused variable warning
    // dependent member template requires 'template'
    t.template bar<double>();
}

The typename keyword disambiguates a dependent type. The template keyword disambiguates a dependent member template. Without these qualifiers the code fails to compile. Experienced developers encounter these errors frequently when writing generic libraries.

#include <print>

struct HasType {
    using type = int;
    template<typename U>
    void bar() { std::println("bar<{}> called", typeid(U).name()); }
};

template<typename T>
void demo(T t) {
    // Dependent type requires 'typename'
    typename T::type *ptr = nullptr;
    (void)ptr; // silence unused
    // Dependent member template requires 'template'
    t.template bar<double>();
}

int main() {
    HasType obj;
    demo(obj);
    std::println("dependent demo ok");
}

A practical illustration is std::vector<T>::iterator. Inside a template that works with an arbitrary container, writing typename C::iterator it is required. The compiler treats iterator as a static member.

Instantiation

The compiler generates concrete code when a template is used. This process is called implicit instantiation. Each distinct set of template arguments triggers a separate instantiation. Implicit instantiation occurs at the point of first use. Errors inside the template body appear only when a particular specialization is needed.

The late-error property has a downside. Because a member is only compiled when instantiated, a typo in an unused branch of a template can escape every build until a caller instantiates that branch. Concepts (chapter 17) address this by checking constraints before instantiation, so mismatches surface at the call rather than deep inside a library.

Explicit instantiation forces the compiler to emit code for a given set of arguments, even if the program never mentions the specialization. This technique can reduce compile time in large builds because the implementation can be compiled once and reused across translation units.

Templates are not free for the build system. Every instantiation produces object code, so a template used with many distinct types can increase binary size and compile time. Explicit instantiation and the extern template declaration control this growth, which matters in large code bases where template-heavy headers are included widely.

extern template class std::vector<int>; // suppress implicit instantiation

A matching template class std::vector<int> line in another translation unit triggers explicit instantiation. The example program demonstrates explicit instantiation and prints the size of a vector to prove that the code was generated.

#include <vector>
#include <print>

// Explicit instantiation of std::vector<int>
template class std::vector<int>;

int main() {
    std::vector<int> v{1,2,3};
    std::println("size = {}", v.size());
    return 0;
}

Explicit instantiation also interacts with the one‑definition rule. The definition must appear in exactly one translation unit. Otherwise the linker reports multiple definition errors. This rule reinforces the importance of a clear build structure.

The examples instantiate at runtime, but templates also compute at compile time. Chapters 18 and 21 use this to perform type‑level computation and static dispatch. Templates can take non‑type parameters, for example an integer size in std::array<T, N>, so the length is known at compile time. Chapter 19 covers this in detail.

Try this

Write a template clamp that limits a value to a closed interval.

template<typename T>
T clamp(T v, T lo, T hi) { /* return lo if v < lo, hi if v > hi, otherwise v */ }

Call it with a value below lo, a value between lo and hi, and a value above hi.

Concepts: the constraint language

Why concepts

Templates allow algorithms to work with any type that satisfies a set of requirements. Before C++20 those requirements were expressed with SFINAE tricks such as std::enable_if and trait metafunctions. The constraints were hidden inside long type expressions, and a failure produced a cascade of template‑instantiation diagnostics that were difficult to read. Concepts replace that style with explicit, named predicates. A concept is part of the function signature. The compiler checks it before it attempts to instantiate the template. If the requirement is not met, the diagnostic cites the concept name and the offending type. The error becomes clear and the intent obvious.

Concepts also act as in‑code documentation: the concept name (e.g., std::integral or std::range) states the precondition next to the declaration, removing the need for separate enable_if blocks. A concept is a compile‑time bool value, usable in if constexpr or static_assert, and it drives overload resolution and in‑body branching. Because the compiler checks the constraint before template instantiation, failures appear at the call site with the concept name and offending type, providing earlier, clearer diagnostics than a static_assert inside the function body.

The requires clause and requires expression

A requires clause follows a template declaration and names one or more concepts that must be satisfied.

template<std::integral T>
void f(T t) {
    std::cout << "integral: " << t << '\n';
}

A requires expression appears inside a template body and describes the operations that must exist for a particular set of types.

template<typename T>
requires requires (T a) { { a + a } -> std::same_as<T>; }
T twice(T a) { return a + a; }

If the predicate evaluates to false, the overload is removed from overload resolution. The following example demonstrates two overloads, one for integral types and one for floating‑point types, and prints which overload ran.

#include <iostream>
#include <type_traits>

// Overload for integral types
template<std::integral T>
void show(T value) {
    (void)value;
    std::cout << "integral overload\n";
}

// Overload for floating‑point types
template<std::floating_point T>
void show(T value) {
    (void)value;
    std::cout << "floating overload\n";
}

int main() {
    show(42);          // integral
    show(3.14);        // floating
    return 0;
}

The requires expression can be combined with logical operators to form richer constraints:

template<typename T>
requires (std::integral<T> && sizeof(T) >= 4)
void process(T value) {
    std::cout << "'integral (sizeof >= 4)'} " << value << '\n';
}

The compiler evaluates the whole Boolean expression before overload resolution, eliminating surprising matches.

The requires clause and the requires expression are distinct tools. The clause (requires std::integral<T>) attaches a constraint to a declaration. The expression (the requires (T a) { ... } block) describes the operations a type must support and can test return types with -> std::same_as<U>. A requires expression can also guard a single member function, so a class template can offer an operation only when the element type supports it.

Defining a concept

A concept is a named predicate. It can be written with the concept keyword and a requires clause that enumerates the required expressions.

#include <iostream>
#include <type_traits>

// Concept that requires addition
template<typename T>
concept Addable = requires (T a, T b) { a + b; };

// Function using the concept
template<Addable T>
T add(T a, T b) {
    return a + b;
}

int main() {
    std::cout << "int add: " << add(2, 3) << std::endl;
    std::cout << "string add: " << add(std::string{"hi"}, std::string{"!"}) << std::endl;
    return 0;
}

The Addable concept checks that the binary + operator is valid for two operands of type T. The accompanying add function uses the concept as a constraint, so any type that models Addable can be added safely. Because the concept name appears directly in the signature, the intent is clear at the call site.

Concepts can also be expressed as Boolean formulas of other concepts. This enables a small vocabulary of high‑level requirements while reusing primitive concepts from <concepts>:

template<typename T>
concept Number = std::integral<T> || std::floating_point<T>;

A library author can build more expressive constraints by layering concepts.

A requires expression enumerates several checks separated by semicolons inside the brace, each a statement that must be valid. Return-type constraints use an arrow: { a + b } -> std::same_as<T> demands that a + b not only compiles but also yields a type convertible to T. Because a concept reduces to a constexpr bool, you can combine concepts with &&, ||, and ! to build richer requirements without writing a new named concept.

The same Boolean nature lets a function branch on a concept with if constexpr, selecting an implementation without writing separate overloads. if constexpr (std::is_pointer_v<T>) { /* pointer path */ } else { /* value path */ } picks the branch at compile time and discards the unused one, so the body need not be valid for every T. This replaces the old tag-dispatch and SFINAE patterns for in-body selection.

Subsumption

When two concepts overlap, the more specific one wins. This rule is called subsumption and mirrors goal ordering in Prolog: the most specific rule is chosen before a more generic one. The compiler builds a partial ordering of candidate functions based on the subsumption relationship of their constraints.

#include <iostream>
#include <type_traits>

// Overload for integral types
template<std::integral T>
void foo(T) {
    std::cout << "integral overload" << std::endl;
}

// Overload for signed integral types (more specific)
template<std::signed_integral T>
void foo(T) {
    std::cout << "signed integral overload" << std::endl;
}

int main() {
    unsigned int u = 1;
    int s = -1;
    foo(u); // should select integral overload
    foo(s); // should select signed integral overload
    return 0;
}

Subsumption removes ambiguity from overload sets. When two constrained overloads both match, the compiler prefers the overload whose constraint subsumes the other. For example, std::signed_integral implies std::integral. Calls with a signed type select the std::signed_integral overload, while unsigned calls select the generic overload. This deterministic ordering eliminates the ambiguous‑overload errors common with SFINAE tricks.

Terse syntax

C++26 allows constraints to appear directly on a parameter type or on a plain auto placeholder.

void show(std::integral auto x) { std::cout << x << '\n'; }
std::integral auto n = 5;

The same pattern works for references, forwarding references, and ranges. The example below uses a constrained variable, a constrained parameter, and a call to std::ranges::sort.

#include <iostream>
#include <vector>
#include <algorithm>
#include <ranges>

// Constrained parameter
void show(std::integral auto n) {
    std::cout << "value: " << n << '\n';
}


int main() {
    std::integral auto n = 7; // constrained variable
    show(n);
    std::vector<int> v = {5, 2, 9, 1, 4};
    std::ranges::sort(v);
    std::cout << "sorted:";
    for (int x : v) std::cout << ' ' << x;
    std::cout << std::endl;
    return 0;
}

These terse declarations keep the constraint next to the entity it describes, improving readability. Because the constraint travels with the parameter, a constrained lambda passed to std::ranges::sort is checked the moment the call is written, not when the algorithm instantiates. They also work inside generic lambdas. This enables concise constraint specifications:

auto cmp = [] (std::totally_ordered auto const& a, std::totally_ordered auto const& b) {
    return a < b;
};

Constraints also attach to return types through the same auto syntax: std::integral auto square(std::integral auto x) { return x * x } constrains both the parameter and the result. Algorithm authors lean on std::predicate and std::relation to describe the callables they accept, so a sorting routine declares std::predicate<std::weak_ordering, T, T> instead of a vague template parameter.

Constrained algorithms and ranges

Standard algorithms already carry concept requirements. std::ranges::sort requires std::sortable. std::ranges::find requires std::range and an std::indirectly_comparable predicate. The compiler checks those concepts before instantiating the algorithm, so a call that does not meet the requirements fails at the call site with a clear diagnostic.

#include <iostream>
#include <vector>
#include <algorithm>
#include <ranges>

int main() {
    std::vector<int> v = {1, 2, 3, 4, 5};
    auto it = std::ranges::find(v, 3);
    if (it != v.end()) {
        std::cout << "found" << std::endl;
    } else {
        std::cout << "not found" << std::endl;
    }
    return 0;
}

The program searches a std::vector<int> for a value. Because the container satisfies std::range, the call compiles. A raw array fails the range concept, and the compiler reports that the concept is not satisfied. Since the concepts appear in the algorithm’s signature, users see the preconditions directly without consulting external documentation or static assertions.

Library concept vocabulary

The header <concepts> provides the core building blocks used throughout the standard library:

  • std::integral: all integral types.
  • std::floating_point: all floating‑point types.
  • std::same_as<T, U>: two types are identical.
  • std::convertible_to<From, To>: implicit conversion is possible.
  • std::invocable<F, Args...>: a callable can be invoked with the given arguments.
  • std::range<R>: a type provides begin and end iterators.

The iterator library adds concepts such as std::input_iterator, std::forward_iterator, std::random_access_iterator, and std::contiguous_iterator. The ranges library refines those with concepts like std::sized_range and std::view. These names become part of the public interface. Reading an algorithm signature tells you exactly which properties the arguments must provide.

Several concepts describe callables rather than types. std::predicate<P, Args...> is true when P can be called on Args... and the result is convertible to bool, which is exactly what std::ranges::find_if requires of its comparator. std::relation and std::strict_weak_order capture the ordering contracts that sorting and set operations depend on. Using them in your own signatures lets the compiler reject a comparator that returns the wrong category before any element is compared.

The std::predicate and std::relation concepts are exactly the constraints behind the algorithms from chapters 12 and 13, so the range pipelines written there were concept-checked whether or not the signatures made it explicit.

Concepts are the modern replacement for the SFINAE machinery that chapter 21 teaches as reading fluency. Where an old signature used std::enable_if in a return type, a new signature writes the concept in the parameter list. The behaviour is equivalent, but the new form is checkable before instantiation and readable at a glance.

Reading concept diagnostics

When a constraint fails, the compiler prints the name of the concept and the type that caused the failure. For example, compiling the following code:

template<std::integral T>
void g(T) {}

g("text");

produces a diagnostic similar to:

error: static assertion failed: constraints not satisfied required from ‘std::integral<char const*>’ note: substitution failed for ‘T = char const*’ The message shows that std::integral was the offending concept and that char const* does not satisfy it. This is far clearer than the cascades of template‑instantiation errors that appeared before concepts.

When user‑defined concepts are involved, the diagnostic includes the failed sub‑requirement. This makes it possible to pinpoint exactly which operation is missing. For instance, if Addable were violated because + returned the wrong type, the compiler points to the { a + b } -> std::same_as<T> clause.

Try this

Define a concept Comparable that requires both a < b and a == b to be valid. Then write a min function constrained by Comparable. Call the function with two int values and with two std::string values.

Compile-time C++

constexpr deeply

constexpr marks a function, variable, or constructor as eligible for constant evaluation. The standard defines a constant expression as an expression that can be evaluated during translation when all operands are themselves constant. When the compiler encounters a call to a constexpr function in such a context, it substitutes the computed value directly into the program.

Since C++14 a constexpr function can contain loops, local variables, and if statements, so its body resembles ordinary code. The constant‑evaluation engine still enforces a sandbox: no I/O, no asm, and no use of the address of a non‑constant object. If a call cannot be evaluated, the compiler either falls back to a runtime call where legal or issues a hard error in a context that requires a constant expression.

Since C++14 a constexpr function can contain loops, local variables, and if statements, so the body looks like ordinary code. The compiler still enforces a constant-evaluation sandbox: no I/O, no asm, and no use of the address of a non-constant object. When a call cannot be evaluated, the compiler either falls back to a runtime call where that is legal, or reports a hard error in a context that demands a constant expression.

The most common pattern is a recursive algorithm that terminates at compile time. The classic example is factorial:

#include <iostream>

constexpr long long factorial(int n) {
    return n <= 1 ? 1 : n * factorial(n - 1);
}

static_assert(factorial(5) == 120, "factorial compile‑time test");

int main() {
    std::cout << "factorial ok\n";
    return 0;
}

The static‑assert in the file forces the compiler to evaluate factorial(5) at compile time. The result of 120 becomes part of the program’s constant pool. The main function prints a short marker so the book’s test harness can verify that the binary linked and executed. This example also demonstrates that constexpr functions can be called from other constexpr contexts, such as template non‑type parameters, std::array sizes, or static_assert conditions.

consteval (immediate functions)

consteval is a stronger guarantee introduced in C++20. An immediate function must be evaluated at translation time. Any attempt to call it where a constant expression is not required is ill‑formed. The compiler therefore rejects the program outright, which produces a diagnostic that points to the offending call site. Immediate functions are ideal for compile‑time utilities that must never appear in the generated binary, such as compile‑time string hashing, type‑level identifiers, or compile‑time parsing of literals.

The following example computes a simple additive hash of a string literal. Because the function is declared consteval, the call hash("abc") is forced into the constant‑evaluation engine. The resulting value is verified with a static_assert. The main function prints a marker that confirms the program compiled successfully.

#include <iostream>
#include <cstddef>

// Very simple compile‑time hash: sum of character codes.
consteval std::size_t hash(const char* str) {
    std::size_t h = 0;
    for (std::size_t i = 0; str[i] != '\0'; ++i) {
        h += static_cast<std::size_t>(str[i]);
    }
    return h;
}

static_assert(hash("abc") == ('a' + 'b' + 'c'), "hash compile‑time test");

int main() {
    std::cout << "hash ok\n";
    return 0;
}

If a programmer later tries to invoke hash with a run‑time string, the compilation fails with a clear message: call to a consteval function is not a constant expression.

Prefer consteval when the function exists only to compute compile-time values, such as a hash or a table generator, because it removes the runtime path entirely and lets the optimizer assume the result is a constant. Prefer constexpr when the same logic also serves runtime inputs, as a parser or a math helper does.

constinit

Static or thread‑local objects with static storage duration are normally zero‑initialized first and then later given their dynamic initializer. This two‑step process can lead to the infamous static‑initialization‑order fiasco when one translation unit accesses a global defined in another before its dynamic initializer runs. The constinit specifier forces the initializer to be a constant expression, guaranteeing that the object is fully initialized before any dynamic initialization begins.

The example below defines a global counter whose value is computed in a constinit variable. The lambda runs at compile time, sums the numbers 0 through 4, and the result 10 becomes the object’s constant initial value. A static_assert validates the result, and the program prints a marker.

#include <iostream>

// Compute a compile‑time sum using a constexpr lambda.
constinit const int global_counter = []constexpr noexcept{
    int x = 0;
    for (int i = 0; i < 5; ++i) x += i;
    return x;
}();

static_assert(global_counter == 10, "constinit compile‑time init");

int main() {
    std::cout << "constinit ok\n";
    return 0;
}

Because global_counter is constinit, the compiler must emit an error if its initializer cannot be evaluated at compile time. This eliminates a whole class of runtime bugs without any additional runtime checks.

constinit differs from a constexpr variable. A constexpr variable is itself constant and can be used in constant expressions. A constinit variable can be non-constant, so it can be mutated later, but its initializer must be constant. Use constinit for mutable globals that must avoid the initialization-order fiasco.

std::is_constant_evaluated

Inside a constexpr function the standard library provides std::is_constant_evaluated(). It returns true when the current evaluation is performed by the constant‑evaluation engine, and false when the function runs at run time. This enables a single implementation to take two distinct paths: a highly optimized compile‑time algorithm and a more flexible run‑time fallback.

The branch example below illustrates this technique. When called inside a static_assert, the compile‑time branch returns 0. When called from main, the run‑time branch prints the word runtime and returns the argument unchanged. The static‑assert confirms the compile‑time result, and the program output shows the run‑time branch execution.

#include <iostream>
#include <type_traits>

constexpr int branch(int x) {
    if (std::is_constant_evaluated()) {
        return 0; // compile‑time path
    } else {
        std::cout << "runtime\n";
        return x;
    }
}

static_assert(branch(42) == 0, "branch compile‑time result");

int main() {
    constexpr int result = branch(7);
    std::cout << "branch result: " << result << "\n";
    return 0;
}

Using std::is_constant_evaluated is a common idiom for providing cheap compile‑time shortcuts without sacrificing generic run‑time behavior. It also avoids the need for separate overloads guarded by if constexpr in user code.

The same test appears inside the standard library. std::vector and std::string check is_constant_evaluated internally to choose an allocation-free path during constant evaluation. Recognizing this pattern helps you understand why a container can be constexpr when a raw allocation cannot be.

Transient constexpr allocation

C++20 lifted the restriction that dynamic allocation cannot appear in a constant expression. The rule states that any allocation performed during constant evaluation is transient: the allocated storage exists only for the duration of the evaluation and is reclaimed automatically when the evaluation finishes. This makes it possible to build containers such as std::vector or std::string inside a constexpr function, as long as the container does not escape the evaluation.

The example builds a std::array via a constexpr helper that fills the elements with squared indices. Although std::array does not allocate dynamically, the pattern demonstrates how a compile‑time algorithm can populate a fixed‑size aggregate. The static_assert checks the final element, and the program prints a marker.

#include <iostream>
#include <array>

constexpr std::array<int,5> make_array() {
    std::array<int,5> a{};
    for (std::size_t i = 0; i < a.size(); ++i) a[i] = static_cast<int>(i * i);
    return a;
}

constexpr auto arr = make_array();
static_assert(arr[4] == 16, "constexpr array element");

int main() {
    std::cout << "array ok\n";
    return 0;
}

On compilers that fully support transient allocation, the same pattern can be written with std::vector:

constexpr std::vector<int> make_vec() {
    std::vector<int> v;
    for (int i = 0; i < 5; ++i) v.push_back(i * i);
    return v; // OK in C++20 and later
}

If the toolchain does not yet implement this feature, the fallback to a fixed‑size container still provides a compile‑time constant data structure.

A practical consequence is that compile-time data tables can be built with ordinary loops and containers, then baked into the binary as constants. This removes both the runtime construction cost and the need to hand-write static initializer lists.

Compile‑time as pure evaluation (Lisp framing)

A constexpr function behaves like a pure term‑rewriter: given inputs it produces an output without observable side effects. The compiler treats it as a mathematical function, can memoize results for identical constant arguments, and folds the value into the program. This mirrors the term‑rewriting semantics of templates introduced in Chapter 16, where the compile‑time engine performs substitution, checks constraints, and yields a result.

When invoked repeatedly with the same constant arguments, the compiler can emit the value once and reuse it, reducing code size and eliminating redundant work. This is analogous to a Lisp interpreter evaluating a pure function at compile time and storing the result for later calls.

The convention flip: static‑assert test suites

Historically, examples printed values and asked the reader to run the program and verify the output manually. With constexpr, the result is known at compile time, so the testing strategy flips to compile‑time verification: a static_assert encodes the expectation directly in the source file. The compiler checks it during translation. A failure stops the build with a clear diagnostic reported by the CI pipeline. The chapter still includes a short std::cout marker so the CTest harness can confirm that the binary linked and executed.

C++26 constexpr frontiers

C++26 expands the constexpr toolbox with several ambitious features that are not yet supported by the Clang 22.1.8 toolchain used for verification.

Not yet deployable. constexpr placement‑new permits constructing an object in a pre‑allocated buffer during constant evaluation. Example syntax:

constexpr char buffer[sizeof(MyType)];
constexpr MyType* p = new (buffer) MyType{ /* args */ };

Not yet deployable. constexpr exception handling makes it possible to throw and catch inside a constant expression. Example syntax:

constexpr int safe_div(int a, int b) {
    if (b == 0) throw std::runtime_error("divide by zero");
    return a / b;
}
static_assert(safe_div(4,2) == 2);

Not yet deployable. User‑generated static_assert messages can incorporate constexpr data, which allows a compile‑time error to report a computed value. Example syntax:

template<int N>
struct prime_checker {
    static constexpr bool value = /* primality test */;
    static_assert(value, "Value " #N " is not prime");
};

These proposals (P2002R2 for placement‑new, P2003R1 for constexpr exceptions, and P2004R0 for message templating) are documented in the C++26 draft but remain unavailable in the current verification environment. The book marks them with a callout so readers understand the future direction without being blocked by the existing toolchain.

Try this

Write a constexpr function is_prime that returns true if its argument is a prime number and false otherwise. Add static_assert checks for a few known values and print a marker from main.

Values as template arguments and user-defined literals

Non‑type template parameters (NTTPs)

Non‑type template parameters turn compile‑time values into part of the type system. By embedding a constant in a template argument the compiler generates a distinct specialization for each value. This enables zero‑overhead dispatch: the generated code contains only the paths required for the specific constant, and any branches that depend on the value disappear after constant folding. Standard library containers such as std::array<T,N> and std::span rely on this technique to expose size information without runtime storage. Compile‑time bounds checking, static‑asserted preconditions, and compile‑time hash tables also become possible when the size or key is an NTTP.

The auto NTTP form, template<auto V>, deduces the argument type. It accepts an integral, pointer, reference, or structural class object. This pattern forwards a value without naming its type, reducing boilerplate. The compiler encodes the value in the mangled name, giving each distinct argument a unique symbol. Using many distinct values can increase binary size, but the gain is eliminating runtime conditionals. It accepts an integral, a pointer, a reference, or a structural class object. Generic utilities frequently use this pattern to forward a value to another template without naming its type, reducing boilerplate and improving readability. The compiler records the value in the mangled name of the instantiation, so each distinct argument yields a unique symbol. This can increase binary size if many values are used, but the trade‑off is often worthwhile for the performance gain of eliminating runtime conditionals.

When an NTTP participates in overload resolution, the compiler prefers a more specialized non‑type argument. This mirrors concept overload resolution and enables tag‑dispatch based on constexpr values. For example, a function template can provide a fast path for a power‑of‑two size by matching template<std::size_t N> requires (N & (N-1)) == 0.

C++26 extends NTTPs to floating‑point values and to class types with non‑trivial constructors, simplifying patterns that rely on encoding a floating value in an integral representation.

Before C++26, a floating‑point NTTP required a workaround: encode the value as a fraction of two integers, or store it in a constexpr static variable and pass a reference. The extension makes the value a direct template argument, so a template can specialise on a literal 1.5 the same way it specialises on an integer. The class-type extension lifts the structural-type restriction that required aggregate initialisation, so a type with a constexpr constructor can now appear as an NTTP even if it has a non-trivial one.

#include <iostream>

// Basic non‑type template parameter: size of an array.
// The size N is a compile‑time constant supplied as a template argument.

template<int N>
struct Arr {
    int data[N]{}; // array of N ints, default‑initialized to zero
};

int main() {
    Arr<5> a; // N = 5
    std::cout << "Arr size: " << sizeof(a) << " bytes" << std::endl;
    return 0;
}

Compile‑time hashing of string literals also relies on NTTPs. A hash function defined as constexpr can accept a fixed_string NTTP and produce a constant hash value. This value can be used as a case label in a switch statement. The result is a zero‑overhead dispatch table for command strings.

An NTTP argument must come from one of a fixed set of categories:

  • Integral and enumeration constants (int, char, enum).
  • Pointers or references to objects with static storage duration, including function pointers.
  • Class‑type values that satisfy the structural type requirements.
  • auto can deduce the type. This allows template<auto V> to accept any of the above categories.

Class‑type NTTPs (C++20)

C++20 extends NTTPs to accept structural types. A structural type is a class or struct whose members are all public, have non‑mutable types, and themselves are structural. The class must not declare any user‑provided constructor, so it can be aggregate‑initialized.

#include <iostream>

// Structural non‑type template parameter.
// The type must be a structural type: all members are public, no mutable, no user‑declared constructor.

struct Config {
    int a;
    double b;
};

// The template takes a Config as a compile‑time value.

template<Config C>
struct UseConfig {
    static constexpr int val_a = C.a;
    static constexpr double val_b = C.b;
};

int main() {
    // Instantiate with a concrete Config value.
    using My = UseConfig<Config{3, 4.5}>;
    static_assert(My::val_a == 3);
    static_assert(My::val_b == 4.5);
    std::cout << "config a=" << My::val_a << " b=" << My::val_b << std::endl;
    return 0;
}

Config satisfies the structural rules: it contains two data members, both of which are fundamental types, and the struct has no constructors. The template UseConfig receives a concrete Config{3,4.5} as a compile‑time constant. Inside the specialization the values appear as constexpr static data members. The static_asserts verify the values during translation, while the program prints a short marker confirming successful runtime execution.

Structural NTTPs enable compile‑time configuration tables, policy objects, and domain‑specific languages without runtime overhead. The structural‑type rule requires all members to be public and non‑mutable, so the compiler can compare two NTTPs for equality during overload resolution and template deduplication. Two aggregates are equal when their members are equal, which keeps hashing and name mangling well defined.

The structural-type rule exists so the compiler can compare two NTTPs for equality during overload resolution and template deduplication. Because every member is public and the class has no user constructor, two aggregate values are equal exactly when their members are equal, which keeps hashing and name mangling well defined. This is what lets a Config{3, 4.5} name one unique specialization.

The fixed_string recipe

Passing a literal string as an NTTP is a common requirement, but the language does not provide a built‑in string NTTP. A small helper class called fixed_string fills this gap. It stores the characters of a literal in a constexpr array and supplies a deduction guide so that a plain string literal can be used directly as a template argument.

#include <iostream>
#include <cstddef>

// Fixed‑string non‑type template parameter.
// Stores the characters of a literal at compile time.

template<std::size_t N>
struct fixed_string {
    char buf[N]{};
    // constexpr constructor copies characters.
    constexpr fixed_string(const char(&s)[N]) {
        for (std::size_t i = 0; i < N; ++i) buf[i] = s[i];
    }
    // constexpr size query.
    constexpr std::size_t size() const { return N; }
};

// Deduction guide from a string literal.
template<std::size_t N>
fixed_string(const char(&)[N]) -> fixed_string<N>;

// Example function taking a fixed_string NTTP.
template<fixed_string S>
void print_fixed() {
    std::cout << "fixed_string: " << S.buf << " (size=" << S.size() << ")" << std::endl;
}

int main() {
    print_fixed<"hello">();
    return 0;
}

fixed_string is a class template parameterised by the literal length N. Its constexpr constructor copies each character from the literal into the internal buffer buf. The deduction guide fixed_string(const char(&)[N]) -> fixed_string<N> allows the user to write print_fixed<"hello">() without explicitly naming the size.

The function print_fixed receives the fixed_string as an NTTP and can access the characters at compile time. The size() member reports the length, and the runtime std::cout prints the stored characters. This technique underlies many compile‑time parsers, compile‑time reflection utilities, and the constexpr‑SQL capstone in Chapter 22.

Because the characters live in the buffer, a consteval parser can walk S.buf at compile time to validate or transform the literal before any runtime code exists. This is the foundation of the compile-time parsing that chapter 22 exploits for its SQL engine: the query string is checked, tokenized, and turned into row types while the program is still being translated.

The structural-type rule is what makes fixed_string legal as an NTTP. Every member is public, the class has no user-declared constructors beyond the constexpr one, and the array of char is a structural type. Two fixed_string values with the same characters produce the same template argument, so the compiler can deduplicate the specialization. This is the same rule that lets a Config{3, 4.5} name one unique specialization, applied to text instead of numbers.

User‑defined literals (UDLs)

A user‑defined literal extends the set of literal suffixes that the language recognises. The compiler looks for a function named operator""_suffix in the same namespace as the call site (or via argument‑dependent lookup). The function can accept an integral, floating‑point, character, or string literal form.

#include <iostream>

struct Distance {
    double value;
    const char* unit;
};

// User‑defined literal for kilometres.
constexpr Distance operator""_km(long double v) {
    return Distance{static_cast<double>(v), "km"};
}

int main() {
    constexpr Distance d = 5.0_km;
    std::cout << "Distance: " << d.value << " " << d.unit << std::endl;
    return 0;
}

The example defines a literal suffix _km that converts a floating‑point literal into a Distance object. The operator""_km is constexpr, so the expression 5.0_km is computed at compile time. The program prints Distance: 5 km, which the test harness verifies with the regular expression Distance: 5 km.

User‑defined literals give a natural, readable syntax for domain‑specific units, bit masks, or compile‑time identifiers. Because they are ordinary functions they can be overloaded, templated, and placed in header files for reuse across translation units.

A UDL has a form for each literal category. Integer literals dispatch to operator""_suffix(unsigned long long), floating-point literals to operator""_suffix(long double), character literals to operator""_suffix(char), and string literals to operator""_suffix(const char*, std::size_t). The _suffix identifier must begin with an underscore. A suffix without a leading underscore is reserved for the implementation. Placing the operator in a namespace and bringing it in with using keeps the literal vocabulary explicit.

UDLs for parsed literals (constexpr)

A more advanced use of UDLs parses the literal characters at compile time. It turns a string literal into a value without any runtime work. The function must be consteval (or constexpr in C++20) and can perform arbitrary compile‑time computation.

#include <iostream>
#include <cstddef>

// consteval user‑defined literal that parses a decimal integer from a string.
// The literal receives the characters and length of the literal.

consteval unsigned int parse_number(const char* str, std::size_t len) {
    unsigned int value = 0;
    for (std::size_t i = 0; i < len; ++i) {
        char c = str[i];
        value = value * 10 + static_cast<unsigned int>(c - '0');
    }
    return value;
}

// UDL suffix _num parses a literal string like "1234"_num into an unsigned int.
consteval unsigned int operator""_num(const char* str, std::size_t len) {
    return parse_number(str, len);
}

int main() {
    constexpr unsigned int v = "1234"_num;
    static_assert(v == 1234);
    std::cout << "Parsed number: " << v << std::endl;
    return 0;
}

operator""_num receives the raw character data and length of the literal. It iterates over the characters, computes the numeric value, and returns it. The parse runs entirely during translation because the literal is a compile-time constant, and the compiler folds the result into the constant v.

In main the literal "1234"_num is parsed into the constant v. A static_assert verifies the result, and the program prints Parsed number: 1234. This pattern is useful for compile‑time parsing of IP addresses, colour codes, or UUID strings.

NTTP callables (C++26)

C++26 adds the ability to pass a callable (a function pointer, lambda, or function object) as a non‑type template argument. The placeholder syntax template<auto F> captures the callable, and the template can invoke it at compile time. The standard library also introduces std::bind_front and std::bind_back, which adapt a callable by pre‑binding or post‑binding arguments, and the resulting binder can be used as an NTTP.

Not yet deployable. NTTP callables and std::bind_front/bind_back are part of the C++26 draft. Current Clang versions lack support, so no compiled example is provided. Readers can experiment with GCC 16 or later once the feature lands.

Units example (NTTP ratios)

Compile‑time unit arithmetic can be expressed using std::ratio, a compile‑time fraction type introduced in Chapter 15. By making a ratio an NTTP a template can encode a conversion factor that the compiler evaluates without runtime cost.

#include <iostream>
#include <ratio>

// Compile‑time ratio representing a speed unit (length / time).

template<int Num, int Den>
struct SpeedRatio {
    static constexpr int num = Num;
    static constexpr int den = Den;
    static void print() {
        std::cout << "speed = " << num << " m/s" << std::endl;
    }
};

int main() {
    SpeedRatio<10,1>::print(); // 10 metres per second
    return 0;
}

SpeedRatio stores a numerator and denominator as template arguments. The print member writes a human‑readable representation. The values are known at compile time, so any arithmetic involving the ratio is folded by the optimizer. The program prints speed = 10 m/s, matching the test harness expectation.

Because the conversion factors are ratios baked into the type, a m/s result and a km/h result cannot be mixed silently. The compiler rejects arithmetic that requires an unknown conversion and accepts only what a dimension analysis allows. This turns a class of unit errors into compile-time diagnostics, the same payoff as the lifetime analysis of chapter 07 but for dimensional correctness. Each std::ratio pair is a distinct type, so meter and second are unrelated unless a template combines them into meter_per_second.

Advanced NTTP patterns

Beyond basic usage, NTTPs can drive compile‑time state machines, generate lookup tables, and enable perfect‑hash dispatch. By encoding a map of string keys to function pointers as a constexpr array of fixed_string/function pointer pairs, a switch on the NTTP can select the handler at compile time, eliminating any runtime map lookup. This approach is useful for command‑line parsers, protocol dispatch, or embedded DSLs where the set of commands is fixed.

The cost is that each distinct NTTP emits a separate specialization, so a table of one hundred string keys produces one hundred instantiations. The technique pays off when the set of commands is fixed and dispatch is hot. For a set that changes, a runtime std::map is simpler. Perfect hashing over fixed_string keys turns a linear scan of command names into a single arithmetic jump, so a hot dispatch loop stays branch-predictable and never performs a runtime string comparison.

Try this

Write a user‑defined literal _hex that parses a hexadecimal string literal into an unsigned integer at compile time and verify the result with a static_assert.

Specialization, overloading, and customization points

This chapter presents the decision rule for choosing concepts versus specialization, then shows the related mechanisms.

Class template specialization

A class template defines a family of types parameterised by one or more template arguments. A full specialization provides a concrete definition for a specific set of arguments. A partial specialization fixes some arguments while leaving others as parameters.

// Primary template: works for any type T
template <typename T>
struct Printer {
    static void print() { std::cout << "primary" << '\n'; }
};

// Partial specialization for pointer types
template <typename T>
struct Printer<T*> {
    static void print() { std::cout << "pointer" << '\n'; }
};

The primary template is selected when the argument list does not match any partial or full specialization. The partial specialization above matches any pointer type T*. The compiler chooses the most specialised viable definition.

#include <iostream>

// Primary template
template <typename T>
struct Printer {
    static void print() {
        std::cout << "primary\n";
    }
};

// Partial specialization for pointer types
template <typename T>
struct Printer<T*> {
    static void print() {
        std::cout << "pointer\n";
    }
};

int main() {
    Printer<int>::print();      // prints "primary"
    Printer<int*>::print();     // prints "pointer"
    return 0;
}

How the compiler decides

  • Matching: the argument list is compared against each specialization’s pattern.
  • Partial ordering: if more than one specialization matches, the compiler ranks them by how many arguments are fixed. The one with the greater number of fixed arguments is preferred.
  • Full specialization: a specialization that fixes all arguments is the ultimate match and overrides any partial specialization.

Because specializations are an out‑of‑band mechanism, they do not participate in overload resolution. They merely replace the primary definition before the compiler instantiates the class template.

A full specialization is written with an empty template parameter list and a concrete argument: template<> struct Printer<int> { ... };. It is the only form that can introduce a definition with a completely different set of members, because it no longer depends on any template parameter. Partial specializations must still match the primary template’s parameter list in shape, so they can vary the pattern but not the member set arbitrarily. In practice most code needs at most one partial specialization and a handful of full ones.

One caution: the primary template must remain well formed even if only specializations are used, because the compiler instantiates the primary in some contexts before consulting specializations. Keep a sensible default body in the primary and treat specializations as refinements.

The decision rule

When you need different behaviour, ask two questions:

  1. Do the types share the same logical interface? If the answer is yes but the implementation varies based on a property (e.g., pointer vs non‑pointer), prefer concepts or if constexpr inside a generic implementation.
  2. Is the representation fundamentally different? If the answer is no but the underlying storage or layout differs (e.g., a raw array versus a std::span), use specialization.

In short, choice is semantic → use concepts. Representation changes → specialise.

Applying the decision rule consistently avoids scattered overloads and specializations. When a concept describes a semantic property, the requirement appears directly in the function signature and the implementation stays in a single generic body. Use specialization only when the type’s layout or representation differs fundamentally, such as a raw array versus a std::span. This discipline simplifies refactoring and provides clearer compile‑time diagnostics.

Variable templates

Variable templates let you define compile‑time constants that depend on a template parameter. They are the value‑side analogue of function templates.

template <typename T>
constexpr T pi = static_cast<T>(3.1415926535897932385L);

static_assert(pi<double> == 3.1415926535897932385);

The standard library uses this pattern extensively, e.g. the _v suffix for trait variables such as std::is_same_v<T, U>.

#include <type_traits>
#include <iostream>

template <typename T>
constexpr T pi = static_cast<T>(3.1415926535897932385L);

int main() {
    static_assert(pi<double> == 3.1415926535897932385);
    std::cout << "pi<double> = " << pi<double> << '\n';
    return 0;
}

Variable templates are instantiated only when ODR‑used, so they impose no runtime cost. They also serve as a natural place for configuration constants such as std::numeric_limits<T>::max(). They compose, e.g., template<typename T> constexpr T half_pi = pi<T> / 2;, allowing the compiler to fold arithmetic at translation time. Because they are constexpr, they can feed static_assert, if constexpr conditions, or non‑type template arguments, bridging value and type worlds. Variable templates also act as value‑side customization points. A library can declare template<typename T> struct traits; and specialize traits<T>::value, while a variable template like template<typename T> inline constexpr bool is_trivially_copyable_v provides the standard‑preferred concise spelling.

if constexpr as in‑body dispatch

Sometimes a single function needs two completely different implementations, but you do not want to write separate overloads or specializations. if constexpr lets you branch at compile time based on a constant expression.

template <typename T>
void show() {
    if constexpr (std::is_pointer_v<T>) {
        std::cout << "pointer" << '\n';
    } else {
        std::cout << "primary" << '\n';
    }
}

int main() { show<int*>(); show<int>(); }

The compiler discards the unreachable branch, so no ill‑formed code can appear in the omitted path. This technique is especially useful when the two paths require different headers or heavy SFINAE tricks.

#include <type_traits>
#include <iostream>

template <typename T>
void show() {
    if constexpr (std::is_pointer_v<T>) {
        std::cout << "pointer" << '\n';
    } else {
        std::cout << "primary" << '\n';
    }
}

int main() {
    show<int*>();
    show<int>();
    return 0;
}

if constexpr replaces tag dispatch. The compiler discards the unreachable branch, so the same show function works for both pointer and non‑pointer arguments without any overload set.

Specialising std::formatter

std::format formats arbitrary types using the formatter customization point. To make a user‑defined type printable, you specialise std::formatter<T> for your type T.

struct Vec2 { int x, y; };

template <> struct std::formatter<Vec2> : std::formatter<std::string> {
    // parse format spec - we ignore it for simplicity
    constexpr auto parse(format_parse_context& ctx) { return ctx.begin(); }
    // format the value
    auto format(const Vec2& v, format_context& ctx) const {
        return std::format_to(ctx.out(), "({},{})", v.x, v.y);
    }
};

int main() {
    Vec2 v{3,4};
    std::cout << std::format("Vec2: {}", v) << '\n';
}

The specialization lives in the same namespace as std (the only exception allowed by the standard). It enables the type to be used with std::format, std::print, and any other formatting facility that forwards to the std::formatter trait.

#include <format>
#include <iostream>

struct Vec2 { int x, y; };

template <> struct std::formatter<Vec2> : std::formatter<std::string> {
    // No format specifiers – ignore the parse context
    constexpr auto parse(std::format_parse_context& ctx) { return ctx.begin(); }
    auto format(const Vec2& v, std::format_context& ctx) const {
        return std::format_to(ctx.out(), "({},{})", v.x, v.y);
    }
};

int main() {
    Vec2 v{3,4};
    std::cout << std::format("Vec2: {}", v) << '\n';
    return 0;
}

Why specialise instead of overloading?

Overloading operator<< works only for stream‑based APIs. std::format follows a type‑centric design: the formatter is a customisation point object (CPO) that the library calls. By providing a specialization, you integrate with the whole formatting ecosystem without pulling in iostreams.

The parse member decodes the part of the format string that follows the colon, and format writes the value. Inheriting from an existing formatter such as std::formatter<std::string> is a shortcut when you want default spec handling. For a type with no natural spec, parse just returns the begin iterator. The specialization must be visible in the namespace of the type or of std, which is why the standard permits specialising std::formatter for user types even though it otherwise forbids adding to std.

Specialising std::hash

std::unordered_map and std::unordered_set hash their keys through std::hash<Key>. A user type has no default, so to use one as a key you specialise std::hash:

struct Key { int id; std::string name; };

template <>
struct std::hash<Key> {
    std::size_t operator()(const Key& k) const noexcept {
        std::size_t h1 = std::hash<int>{}(k.id);
        std::size_t h2 = std::hash<std::string>{}(k.name);
        return h1 ^ (h2 << 1);
    }
};

The specialization must be a complete type with an operator() returning std::size_t. Combine the hashes of the members with a shift and XOR so the result depends on both fields. Equality must stay consistent with the hash: two keys that compare equal must hash equal, or the container misbehaves. The same pattern powers the customization that std::unordered_map relies on for every key type.

Because the hash and the equality operator must agree, define them together. If you change the equality later, update the hash in the same change or lookups return wrong results silently. The standard library also lets you supply a custom hasher or comparator as template arguments to unordered_map, so a bespoke type can avoid touching std::hash entirely when its needs are unusual.

ADL and hidden friends

Argument-dependent lookup (ADL) finds functions in the namespaces of their arguments. When a == b is written, the compiler looks for operator== not only in the enclosing scope but also in the namespaces of a and b. A hidden friend is an operator== defined inside the class body. It is reachable only through ADL, so it never pollutes the global namespace and cannot be called without the right argument types.

struct Point {
    int x, y;
    friend bool operator==(const Point& a, const Point& b) = default;
};

Because the friend is defined inline, it is a hidden friend and ADL is the only mechanism that finds it. This keeps the operator close to the type, prevents accidental calls on unrelated types, and gives the best diagnostics: a call with a mismatched operand names Point rather than surfacing an unrelated overload. The same reasoning applies to any operator used only with its own operand types.

To summarize, the library can combine concepts, specializations, variable templates, and customization point objects to provide a clear, layered customization strategy that keeps generic code simple and concrete overrides focused.

Try this

Create a small struct Color { uint8_t r,g,b } and write a std::formatter<Color> that formats the colour as a hex string #RRGGBB. Verify the output with std::format.

Reading the old magic: SFINAE, type_traits, and friends

Why you must read old TMP

The C++ ecosystem did not jump from C++98 straight to concepts. Decades of libraries, such as the Standard Library, Boost, Eigen, and countless in-house frameworks, were written using SFINAE, std::enable_if, <type_traits> utilities, tag dispatch, and the Curiously Recurring Template Pattern (CRTP). When you encounter a templated overload that appears cryptic, you must not rewrite it immediately. First, decode the intent:

  1. Identify the constraint expressed by the enable-if or trait.
  2. Translate that constraint into a requires clause or a named concept.
  3. Verify that the modern form selects the same overload.

This disciplined approach preserves the original algorithmic intent while allowing you to modernise gradually. Moreover, many code reviewers still expect you to read SFINAE-heavy code, because refactoring large libraries without breaking ABI is risky. By mastering the legacy patterns you become a bridge between the old and the new.

SFINAE and enable_if are not a separate language. They are template deduction rules pushed to their limits. Reading them as such, rather than as magic, makes the modern forms obvious. When you see std::enable_if_t<...> = 0 in a signature, translate it in your head to the pre-concepts spelling of a requires clause, and the intent snaps into focus.


std::enable_if positions

std::enable_if can appear in three classic locations:

  • Return type: the function’s result type is conditionally enabled.
  • Default template parameter: the template parameter list carries a hidden enable_if that activates the overload.
  • Parameter type: the function parameter itself is wrapped in enable_if, often for std::string arguments.

Immediate-context rule: a substitution failure in the part of a declaration that participates in overload resolution removes the candidate without a diagnostic. This is the basis of SFINAE (Substitution Failure Is Not An Error). It lets enable_if expressions in return types, default template parameters, or parameter types silently disable overloads.

Below is a single file that demonstrates all three signatures. The program prints a message that identifies the selected overload. The modern rewrite, shown as comments, uses concepts to achieve the same overload resolution.

// ch21_enable_if_decode.cpp
#include <iostream>
#include <type_traits>
#include <string>

// Overload using enable_if in return type (integral types)
template <typename T>
std::enable_if_t<std::is_integral_v<T>, std::string> foo(T) {
    std::cout << "int overload" << std::endl;
    return "int";
}

// Overload using enable_if as a default template parameter (floating point types)
template <typename T, std::enable_if_t<std::is_floating_point_v<T>, int> = 0>
std::string foo(T) {
    std::cout << "float overload" << std::endl;
    return "float";
}

// Overload using enable_if in a parameter type (std::string)
template <typename T>
std::enable_if_t<std::is_same_v<T, std::string>, std::string> foo(T const& value) {
    std::cout << "string overload" << std::endl;
    return value;
}

int main() {
    foo(42);
    foo(3.14);
    foo(std::string{"hello"});
    return 0;
}
// Modern rewrite (concept-based)
template <typename T>
requires std::integral<T>
std::string foo(T) { std::cout << "int overload" << std::endl; return "int"; }

template <typename T>
requires std::floating_point<T>
std::string foo(T) { std::cout << "float overload" << std::endl; return "float"; }

template <typename T>
requires std::same_as<T, std::string>
std::string foo(T const& v) { std::cout << "string overload" << std::endl; return v; }

Running the legacy example yields:

int overload
float overload
string overload

Each overload is selected exactly as the concept-based version is, demonstrating a mechanical one-to-one mapping.

template <typename T, typename = std::void_t<decltype(std::declval<T>().size())>>
void f(T const&) { std::cout << "has size" << std::endl; }

template <typename T>
void f(T const&) { std::cout << "no size" << std::endl; }

If T has a static member size(), the first overload is viable. Otherwise the substitution fails, the candidate is discarded, and the second overload wins. This behaviour underlies the enable_if patterns shown earlier. The enable_if expression lives in the immediate context, so a non-matching type eliminates the overload.


The <type_traits> vocabulary

The <type_traits> header provides a small declarative language that predates concepts. The most common utilities are:

  • std::is_same_v<T,U>: true if T and U denote the same type.
  • std::is_integral_v<T>: true for integral types.
  • std::is_floating_point_v<T>: true for floating-point types.
  • std::conditional_t<C,T,F>: selects T if C is true, otherwise F.
  • std::void_t<...>: used to build detection idioms.

The classic detection idiom uses void_t to test whether a particular expression is well-formed. The example below shows both the legacy and the modern approach.

// ch21_void_t_has_size.cpp
#include <iostream>
#include <type_traits>

// --- Classic detection using std::void_t ---
struct WithSize {
    static constexpr std::size_t size() { return 5; }
};
struct WithoutSize {};

template <typename, typename = void>
struct has_size : std::false_type {};

template <typename T>
struct has_size<T, std::void_t<decltype(T::size())>> : std::true_type {};

// --- Modern detection using a requires expression (concept) ---
template <typename T>
concept HasSize = requires { T::size(); };

int main() {
    std::cout << "has_size<WithSize>::value = " << has_size<WithSize>::value << "\n";
    std::cout << "has_size<WithoutSize>::value = " << has_size<WithoutSize>::value << "\n";
    static_assert(HasSize<WithSize>, "WithSize satisfies HasSize");
    static_assert(!HasSize<WithoutSize>, "WithoutSize does not satisfy HasSize");
    return 0;
}
// Modern rewrite using a requires expression (concept)
template <typename T>
concept HasSize = requires { T::size(); };

Both versions print the same truth values and compile-time static_asserts, confirming that the concept captures the exact same property as the void_t detector.


Tag dispatch

Tag dispatch separates overloads by passing an artificial tag type as an extra argument. The tag is selected by a constexpr function that examines type traits. The pattern is useful when you need a fast, binary-size friendly dispatch that does not rely on SFINAE.

// ch21_tag_dispatch.cpp
#include <iostream>
#include <type_traits>

// Tag types
struct IntegralTag {};
struct FloatingTag {};

// Primary overload – catches any type via tag dispatch
template <typename T>
void print_type(T, IntegralTag) {
    std::cout << "integral" << std::endl;
}

template <typename T>
void print_type(T, FloatingTag) {
    std::cout << "floating" << std::endl;
}

// Helper to select tag based on type traits
template <typename T>
constexpr auto select_tag() {
    if constexpr (std::is_integral_v<T>) {
        return IntegralTag{};
    } else if constexpr (std::is_floating_point_v<T>) {
        return FloatingTag{};
    } else {
        static_assert(!std::is_same_v<T, T>, "Unsupported type");
    }
}

int main() {
    print_type(42, select_tag<int>());
    print_type(3.14, select_tag<double>());
    return 0;
}

The modern replacement uses if constexpr directly inside the function body, eliminating the need for separate tag types:

template <typename T>
void print_type(T) {
    if constexpr (std::is_integral_v<T>) {
        std::cout << "integral" << std::endl;
    } else if constexpr (std::is_floating_point_v<T>) {
        std::cout << "floating" << std::endl;
    } else {
        static_assert(!std::is_same_v<T,T>, "Unsupported type"):
    }
}

Both implementations produce identical runtime output for the test cases. The tag types are cheap, because an empty struct carries no data and the compiler folds the dispatch away entirely, which is why the pattern remains common in performance-sensitive code.


CRTP (Curiously Recurring Template Pattern)

CRTP is a form of static polymorphism where a base class template takes the derived class as a parameter. The base can call functions that the derived implements. It uses static_cast to do so. This pattern appears in many older libraries (e.g., Boost.Fusion, Eigen) and remains useful for compile-time interfaces.

// ch21_crtp_shape.cpp
#include <iostream>
#include <type_traits>

// CRTP base that provides an interface for area()
template <typename Derived>
struct ShapeBase {
    // Calls the derived implementation via static_cast
    double area() const {
        return static_cast<const Derived*>(this)->area_impl();
    }
};

// Circle implementation using CRTP
struct Circle : ShapeBase<Circle> {
    double radius;
    explicit Circle(double r) : radius(r) {}
    double area_impl() const { return 3.1415926535 * radius * radius; }
};

// Rectangle implementation using CRTP
struct Rectangle : ShapeBase<Rectangle> {
    double width, height;
    Rectangle(double w, double h) : width(w), height(h) {}
    double area_impl() const { return width * height; }
};

int main() {
    Circle c(2.0);
    Rectangle r(3.0, 4.0);
    std::cout << "Circle area: " << c.area() << "\n";
    std::cout << "Rectangle area: " << r.area() << "\n";
    return 0;
}

A concept-based modern alternative expresses the same requirement with a requires clause on a free function or a generic algorithm, avoiding inheritance entirely:

template <typename Shape>
concept ShapeLike = requires(const Shape& s) { s.area_impl(): }:

template <ShapeLike S>
double area(const S& s) { return s.area_impl(): }

The concept-based version provides the same static guarantee (area_impl exists) without the boilerplate of a CRTP base class.

The CRTP also lets a base provide default implementations that call into the derived type, which is how many mixin-style helpers share code without virtual dispatch. Because the call is resolved at compile time, it is inlined. This gives the performance of hand-written code with the reuse of a shared base.


Translation table

The following table summarises the mechanical rewrite from legacy constructs to their modern equivalents. Each row represents a direct replacement that preserves behaviour while simplifying the code.

Legacy patternModern equivalent
std::enable_if in return type, default template parameter, or parameter typerequires clause or named concept
SFINAE (substitution failure)Concept subsumption during overload resolution
Detection idiom with std::void_trequires { expr: } (requires expression)
Tag dispatch (tag type + overload)if constexpr inside a single overload
CRTP static polymorphismConcept-constrained free function or algorithm

Practical migration workflow

When converting legacy TMP code, follow a systematic process. First, locate the overload set that uses enable_if or a detection idiom. Next, write an equivalent requires clause that captures the same logical condition. Then replace the multiple overloads with a single constrained function if possible. After editing, compile the program and verify that the output matches the original. Finally, run the book’s test harness to confirm that the EXPECT strings still pass.

Apply the same methodology to tag dispatch. Identify the tag types and the helper select_tag function. Replace them with an if constexpr chain that performs the same trait checks. Remove the tag definitions and the extra overloads. Compile and test to ensure identical behavior.

For CRTP bases, extract the required interface into a concept. Write free functions that operate on any type satisfying the concept. Remove the CRTP inheritance and replace it with calls to the free functions. Verify that the program still prints the expected results.

By iterating through each pattern in the translation table, you gradually modernize the codebase while maintaining functional parity. This approach reduces technical debt and improves readability for future maintainers.

Try this

Translate the following three-overload enable_if set into a single concept-ordered function family. The original overloads handle (a) integral types, (b) floating-point types, and (c) all other types. After you rewrite, add a static_assert that verifies the same overload is selected for int, double, and std::string.

// Original legacy code (do not modify)
template <typename T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
void process(T) { std::cout << "int" << std::endl; }

template <typename T, std::enable_if_t<std::is_floating_point_v<T>, long> = 0>
void process(T) { std::cout << "float" << std::endl; }

template <typename T, std::enable_if_t<!std::is_arithmetic_v<T>, void*> = nullptr>
void process(T) { std::cout << "other" << std::endl; }

Write the concept-based version and the static_asserts. No solution is provided.


The translation of legacy patterns into concepts is not merely a stylistic upgrade: it has practical impacts on compilation speed, error diagnostics, and maintainability. Modern compilers can evaluate requires expressions early, pruning invalid overloads before template instantiation proceeds, which often reduces template recursion depth and improves compile-time feedback. Error messages tied to a concept name point directly at the violated constraint, a stark contrast to the often-cryptic cascade of errors produced by deep SFINAE failures.

Furthermore, concepts become part of the public interface. When a library author publishes a concept, downstream users can read the requirement in the function signature without hunting through implementation details. This self-documenting quality aligns with the book’s overarching philosophy: code must convey intent as clearly as possible.

When converting existing code, adopt a staged approach:

  1. Identify a SFINAE-enabled overload.
  2. Write an equivalent requires clause, preserving the logical condition.
  3. Replace the overload set with a single constrained function if possible.
  4. Run the test suite to confirm behaviour remains identical.
  5. Refactor additional legacy patterns (type-trait checks, tag dispatch, CRTP) following the same mechanical mapping.

By iterating through the table above, you systematically reduce technical debt, modernise the codebase, and equip future contributors with clearer abstractions. The chapter’s examples demonstrate each step. They give you concrete, compile-tested references for every transformation.

The examples in this chapter are compiled and tested with the book_example macro. The EXPECT strings in the CMake registration verify that each program prints the expected identifiers.

Capstone: a constexpr SQL in C++26

Why a constexpr EDSL

A domain‑specific language (DSL) that runs at compile time gives the same safety as a Lisp macro system but without textual substitution. The query is parsed, type‑checked, and turned into data while the compiler translates the translation unit, so a malformed query produces a compilation error instead of a runtime failure. This idea follows the term‑rewriting view of templates (chapter 16) and the compile‑time execution model (chapter 18).

The lineage is clear. Deane and Turner demonstrated a constexpr JSON parser that turned a string literal into a constexpr data structure. CTParser and the CTRE library later showed how regular‑expression‑based parsers can live in consteval functions. More recently, mkitzan’s constexpr‑sql project combined those ideas to build a tiny SQL‑like EDSL. The capstone brings those ingredients together in a single, self‑contained example.

A constexpr EDSL is not a new parser library bolted onto the build. It uses the same term‑rewriting machinery as the template system applied to a string literal. The query becomes compiler data, not a runtime command, so the compiler checks every branch as it evaluates. A compile‑time failure is the cheapest failure because it occurs before the binary exists and names the exact erroneous query text. Consequently, the compiler acts as the test runner and the query literal serves as the fixture, eliminating the need for a separate runtime test harness.

A compile-time failure is the cheapest kind of failure, because it happens before the binary exists and names the exact query text that is wrong. A runtime test needs a runner, an assertion library, and a suite to maintain. Here the compiler is the runner and the query literal is the fixture.

The building blocks

The capstone re‑uses exactly the mechanisms introduced earlier:

  • fixed_string NTTP (chapter 19) carries the query literal as a non‑type template parameter.
  • consteval (chapter 18) forces the parser to run during translation.
  • constexpr containers: std::array and std::vector with transient allocation hold tables and query results entirely at compile time.
  • if constexpr selects between the supported comparison operators without generating unreachable code.

No other library is required. The whole program lives in a single source file, and each piece (fixed_string, consteval, and constexpr containers) is a plain, non‑exotic component whose combination yields a real, self‑checking domain language without any runtime component.

The schema as aggregates

A table row is a plain aggregate struct, exactly as in chapter 3:

struct Row {
    int id;
    int value;
};

A table is a constexpr array of those rows:

constexpr Row table[] = {
    {1, 10},
    {2, 20},
    {3, 30},
};

Column names are represented by fixed_string NTTPs. The parser extracts the name from the query text, stores it in a constexpr structure, and later matches it against the members of Row.

Keeping the schema as aggregates is deliberate. Because Row is an aggregate and the table is a constexpr array, the compiler can fold the whole table into a constant at translation time, with no allocation, no constructor, and no pointer indirection at runtime. The same reasoning that made std::array the right container in chapter 11 applies here with extra force.

The aggregate discipline also pays off when the schema grows. Adding a column is a one-line change to Row, and the parser and evaluator read the members by name through fixed_string, so no extra wiring is needed for the new field.

Parsing the query at compile time

The parser is a tiny recursive‑descent engine written as a consteval function. It operates on the character buffer of a fixed_string and produces a Query object.

template<std::size_t N>
struct fixed_string {
    char data[N + 1]{};
    constexpr fixed_string(const char (&s)[N + 1]) {
        for (std::size_t i = 0; i < N; ++i) data[i] = s[i];
        data[N] = '\0';
    }
    constexpr const char* c_str() const { return data; }
};

struct Query {
    fixed_string<16> select; // column after SELECT
    fixed_string<16> where;  // column after WHERE
    char op;                // comparison operator, currently '>' only
    int rhs;                // right‑hand side constant
};

// Very small helper that skips whitespace
consteval const char* skip_ws(const char* p) {
    while (*p == ' ' || *p == '\t' || *p == '\n') ++p;
    return p;
}

// Reads an identifier up to the next whitespace or delimiter
consteval std::size_t read_ident(const char* p, char* out) {
    std::size_t i = 0;
    while (*p && *p != ' ' && *p != '\t' && *p != '\n' && *p != ',' && *p != ';') {
        out[i++] = *p++;
    }
    out[i] = '\0';
    return i;
}

consteval int read_number(const char* p, const char*& end) {
    int value = 0;
    while (*p >= '0' && *p <= '9') {
        value = value * 10 + (*p - '0');
        ++p;
    }
    end = p;
    return value;
}

// Compile‑time parser for the supported subset
consteval Query parse_query(const char* qs) {
    const char* p = qs;
    p = skip_ws(p);
    // SELECT keyword (assumed present)
    p += 6; // skip "SELECT"
    p = skip_ws(p);
    char sel[17]{}; read_ident(p, sel);
    p += std::char_traits<char>::length(sel);
    p = skip_ws(p);
    // FROM keyword (ignored, table name is fixed)
    p += 4; // skip "FROM"
    p = skip_ws(p);
    while (*p && *p != ' ') ++p; // skip table identifier
    p = skip_ws(p);
    // WHERE keyword
    p += 5; // skip "WHERE"
    p = skip_ws(p);
    char wh[17]{}; read_ident(p, wh);
    p += std::char_traits<char>::length(wh);
    p = skip_ws(p);
    // operator – only '>' is accepted
    char op = *p;
    ++p; p = skip_ws(p);
    const char* num_end;
    int rhs = read_number(p, num_end);
    Query q{};
    q.select = fixed_string<16>{sel};
    q.where  = fixed_string<16>{wh};
    q.op = op;
    q.rhs = rhs;
    return q;
}

If the input deviates from the supported grammar, a static_assert aborts compilation with a clear diagnostic.

The parser is deliberately simple: it walks the buffer, skips whitespace, reads a column name, skips the fixed keywords, and parses one number. There is no token vector and no abstract syntax tree. Each keyword offset is hard-coded because the grammar fixes the order, so the scanner never needs lookahead. This is the minimum structure that still shows the shape of a real recursive-descent parser.

Producing the result rows

The evaluator walks the constexpr table, applies the predicate, and writes matching values into a fixed‑size std::array. The size of the array is the maximum possible number of rows. The unused slots remain zero.

consteval std::array<int, 3> eval(const Query& q) {
    std::array<int, 3> out{{0, 0, 0}};
    std::size_t idx = 0;
    for (const Row& r : table) {
        int lhs = (std::string_view(q.where.c_str()) == "value") ? r.value : 0;
        bool ok = false;
        if (q.op == '>') ok = lhs > q.rhs;
        if (ok) {
            out[idx++] = (std::string_view(q.select.c_str()) == "id") ? r.id : r.value;
        }
    }
    return out;
}

Returning a fixed-size array is a deliberate simplification. A std::vector is more natural, but a fixed array makes the compile-time result easier to reason about and avoids relying on the details of transient allocation. The WHERE column is looked up by name at compile time with a string comparison against the members.

The static_assert test suite

The capstone follows the chapter‑18 convention of using static_assert as the only test harness. The main function prints a short marker so that the CTest PASS_REGULAR_EXPRESSION matches something.

int main() {
    // The query is parsed at compile time; any syntax error aborts compilation.
    constexpr auto q = parse_query("SELECT id FROM table WHERE value > 15");
    static_assert(std::string_view(q.select.c_str()) == "id", "wrong SELECT column");
    static_assert(std::string_view(q.where.c_str()) == "value", "wrong WHERE column");
    static_assert(q.rhs == 15, "wrong constant");
    constexpr auto result = eval(q);
    static_assert(result[0] == 2, "first matching id");
    static_assert(result[1] == 3, "second matching id");
    std::cout << "constexpr‑SQL capstone OK" << '\n';
}

If any of the static_asserts fail, compilation stops and the programmer receives the message. The binary itself prints only a marker. The real verification lives in the compile‑time checks.

Compiling this file is the test. If any assertion fails, the build stops before a single instruction is executed, and the diagnostic names the failed predicate and the offending constant. There is nothing to run, nothing to instrument, and no chance of a false green from a test that forgot to check. This is the chapter-18 convention applied to a whole program.

Where the capstone stops

The implementation deliberately covers a narrow subset of SQL: a single SELECT column, a single table name (hard‑coded), and a single numeric column in the WHERE clause with the > operator. Extending the parser to recognise additional comparison operators, logical conjunctions, multiple columns, or joins calls for a larger grammar, more token types, and a recursive-descent engine that can build an abstract syntax tree. The same compile-time techniques, namely consteval functions, if constexpr dispatch, and constexpr containers, apply, but the code size grows quickly.

None of the omitted features is impossible. Each is simply more code. A <= operator is one more branch. String columns need a length-aware comparison. Multiple columns need the parser to build a small list of names. The point of the subset is to show the technique works and to keep the example short enough to read in one pass.

Static reflection preview (PROSE ONLY)

C++26 defines a static‑reflection proposal (P2996) that introduces the ^^T operator, std::meta::info, and splicing syntax. In theory those facilities can replace the hand‑written tokenizer with a compile‑time inspection of the query string literal, and they can generate the Query structure automatically.

Not yet deployable.

The feature currently compiles only with GCC 16 (partial support) and is absent from Clang 22.1.8, which the book’s toolchain pins. The capstone therefore implements the parser manually.

When reflection lands, much of this machinery collapses. A compiler-provided ^^T splice can walk the query literal and build the Query directly, and a future library can generate the evaluator from the schema. Until then, the manual parser is the honest, portable baseline that every compiler supports.

Try this

Extend the parser to support a second comparison operator (<). Add a static_assert that a query using < selects the correct rows.

Headers, #include, and organizing programs

The build model everyone uses

Modern C++ programs are built from many translation units. A translation unit is the result of a single source file after the preprocessor has run. The compiler translates each unit independently to an object file and a linker later combines all object files into an executable or a library.

The preprocessor step is a literal text substitution. When the source contains a line such as

#include "vec2.hpp"

the preprocessor opens vec2.hpp, copies every line into the source, and then continues scanning the combined text. No compilation happens before the whole file has been assembled. All macro expansion, conditional compilation, and header inclusion happen at this stage. The resulting file is what the compiler sees as one translation unit.

Because the process is entirely textual, a header can be included many times, from many different source files, and even through different relative paths. If a header is pulled in twice the same name appears twice in the same translation unit, which is illegal unless the header is protected.

Modules, introduced in the next chapter, provide a different mechanism that bypasses textual pasting. For now the dominant reality is the include-paste model described above.

Separate compilation has a practical payoff. When one source file changes, only that unit recompiles, and the linker recombines it with the unchanged object files. This is why a large project rebuilds quickly after a single edit instead of recompiling everything. The compiler sees each unit in isolation, so it learns names from headers.

Declarations versus definitions

A declaration introduces a name to the compiler. A definition supplies the complete entity that the name represents.

Typical examples:

// Declaration – tells the compiler that a function exists.
double length(const Vec2& v);

// Definition – provides the body that implements the function.
double length(const Vec2& v) { return std::sqrt(v.x*v.x + v.y*v.y); }

Headers normally contain declarations. Source files (.cpp) contain the matching definitions.

To call a function, the compiler needs only its declaration: the name, the parameter types, and the return type. The body can live in a different translation unit, compiled separately and found by the linker. For a class, however, the full definition must be visible wherever the class is used by value or its members are accessed, which is why class definitions live in headers.

Templates and inline functions are an exception: the definition must be visible to every translation unit that uses them, so the definition itself lives in a header. The same rule applies to constexpr variables: their definition is required in each unit that odr‑uses the variable.

Include guards and #pragma once

If the same header is included twice, the preprocessor copies the same declarations and definitions twice, producing duplicate symbols and compile‑time errors. Two mechanisms prevent this duplication.

Classical include guard

#ifndef VEC2_HPP          // If VEC2_HPP not defined …
#define VEC2_HPP          // … define it and process the file.

... header contents …

#endif // VEC2_HPP        // End of guarded region.

The first inclusion defines the macro. Subsequent inclusions see the macro already defined and skip the whole file.

#pragma once

Many compilers support the single directive

#pragma once

placed at the top of a header it guarantees the file is processed at most once per translation unit. The effect is identical to the classic guard but requires fewer lines and avoids accidental macro name clashes.

Both forms are accepted by the book’s build. The examples show the guard form for portability.

The One Definition Rule

The One Definition Rule (ODR) states that every non‑inline function, variable, class, or template specialization must have exactly one definition in the entire program.

If two translation units each contain a definition of the same non‑inline function, the linker reports a multiple‑definition error.

Headers are the root cause of many ODR violations. A header that contains a full definition of a normal function will be copied into every translation unit that includes it, creating multiple definitions. That is why most functions belong in source files, while only declarations appear in headers.

The ODR also applies to types: two different definitions of a class with the same name break the rule, even if the definitions are textually identical. The linker cannot merge class definitions. The program must contain a single authoritative definition.

A type-level ODR violation is the subtlest: differing class layouts across units cause memory corruption, so define each class once in a header and include it everywhere. Likewise, placing a non‑inline function definition in a header creates multiple definitions. Move the definition to a source file and keep only the declaration in the header.

inline as the ODR valve

The inline specifier deliberately relaxes the ODR for functions and variables.

  • an inline function can be defined in any number of translation units, provided each definition is identical after preprocessing.
  • an inline variable (C++17 onward) follows the same rule.

When a header defines an inline function or variable, each translation unit that includes the header gets its own copy. The linker discards the duplicates and keeps a single entity.

Because the definitions are required to be identical, the compiler can safely replace a call with the function body (inline expansion) or keep a single out‑of‑line copy if necessary.

The book uses inline constexpr double pi = 3.141592653589793;. The definition must be identical in every translation unit, otherwise behavior is undefined.

Linkage and anonymous namespaces

Linkage determines whether a name is visible across translation units.

Linkage typeVisibility
externalvisible to the linker. The name can be used from any translation unit.
internalvisible only inside the translation unit where it is defined.

The static keyword on a namespace‑scope variable gives it internal linkage (deprecated for functions, but still valid). A const variable at namespace scope also has internal linkage unless explicitly marked extern.

An anonymous namespace provides internal linkage for every name declared inside it:

namespace {
    int helper() { return 42; }   // internal linkage
}

The compiler assigns a unique mangled name to each translation unit, guaranteeing that the helper does not clash with a helper in another file. This technique is useful for implementation‑detail functions that must not appear in the public symbol table.

The choice between internal and external linkage is an interface decision. Names with external linkage are part of the program’s symbol table and can collide with other units. Names with internal linkage are private to their unit, so two units can each define a detail helper with the same name without conflict. This is what makes anonymous namespaces the standard way to hide implementation helpers.

The multi‑file example

The following three files implement a tiny 2‑D vector library. The header declares the type and its interface, the source file defines the functions, and main.cpp uses the library.

#ifndef VEC2_HPP
#define VEC2_HPP

#include <cmath>

// Inline constant – visible to every translation unit.
inline constexpr double pi = 3.14159265358979323846;

struct Vec2 {
    double x{};
    double y{};

    // Defaulted constructors.
    Vec2() = default;
    Vec2(double x_, double y_);

    // Returns the squared length – useful for comparisons.
    double length_sq() const;

    // Returns the Euclidean length.
    double length() const;

    // Adds another vector to this one.
    void add(const Vec2& other);
};

#endif // VEC2_HPP
#include "vec2.hpp"

Vec2::Vec2(double x_, double y_) : x(x_), y(y_) {}

double Vec2::length_sq() const {
    return x * x + y * y;
}

double Vec2::length() const {
    return std::sqrt(length_sq());
}

void Vec2::add(const Vec2& other) {
    x += other.x;
    y += other.y;
}
#include "vec2.hpp"
#include <iostream>

int main() {
    Vec2 v(3.0, 4.0);
    std::cout << "length " << v.length() << '\n';
    return 0;
}

The directory also contains a CMakeLists.txt that builds the program as a raw executable and registers a test that checks the printed output:

add_executable(ch23_two_files main.cpp vec2.cpp)
add_test(NAME ch23_two_files_run COMMAND ch23_two_files)
set_tests_properties(ch23_two_files_run PROPERTIES
    PASS_REGULAR_EXPRESSION "length 5"
)

The build command that a reader runs from the repository root is

cmake --preset dev -S . -B build && cmake --build build --target ch23_two_files

Running the test with ctest --test-dir build confirms that the program prints the expected length of a 3‑4‑5 right triangle.

Compiling the three files separately is instructive. vec2.cpp compiles vec2.hpp, so any mismatch between declarations and definitions fails here. main.cpp compiles vec2.hpp again, and the linker merges the two object files. The header is therefore compiled twice, once per unit, which is exactly why it must be self-contained and guarded, and why editing it forces every including unit to rebuild.

Header hygiene

A self‑contained header compiles on its own. That means it includes every header it needs, and nothing else.

Never rely on a transitive include from another header. If vec2.hpp needs <cmath> it must include it directly, even if another header already includes <cmath>. This prevents surprising compile errors when the header is used elsewhere.

Include what you use is the guiding principle.

If a header only uses a forward declaration of a class, it must forward‑declare rather than include the full definition. This reduces compile‑time dependencies and avoids unnecessary recompilation when unrelated headers change.

The example library follows these rules:

  • vec2.hpp includes only <cmath> because the header needs the std::sqrt declaration for the inline length() definition.
  • the source file vec2.cpp includes the same header to ensure the declarations match the definitions.
  • main.cpp includes only vec2.hpp and the standard <iostream> for output.

Every header must be compiled in isolation to detect missing includes, and forward declarations replace full includes when only pointers or references are used.

Try this

Split the expression‑tree variant example from chapter 3 into two files: a header that declares the Expr type and its visitor, and a source file that defines the evaluation function. Keep the program’s behaviour unchanged and make sure the build still passes the existing test.

Modules, the standard’s direction

What modules fix over #include

The traditional include mechanism copies the text of a header file into each translation unit that names it. It also copies every macro definition that appears before the include directive. Because each translation unit receives its own copy of the header, the compiler cannot share work between units, which leads to long compile times for large projects.

A macro defined in one header can change the meaning of code that includes a later header. Thus, the order of header inclusion influences the program. Errors that originate inside a header are reported at the line that performed the include, making it hard to locate the source of the problem.

Modules replace this textual inclusion with a compiled interface. A module’s interface is built once. This build produces a binary module interface (BMI) file. The BMI contains only the declarations that the author chooses to export. Macros are not part of the BMI, so a macro defined in one translation unit cannot affect a module that imports it. Errors that arise inside a module are reported inside the module file itself, giving a clear location. The build system can reuse the BMI for every importer, which reduces compile time dramatically for projects that import the same module many times. In short, modules give isolation, order independence, and faster incremental builds.

import std

The C++ standard library can be treated as a single module. The statement

import std;

brings every name from the library into the program. No header file needs to be included. The compiler reads the BMI for the standard library module instead of opening dozens of header files. Unfortunately the pinned toolchain (Clang 22.1.8) does not recognise the import keyword, so a file that contains the line above fails to compile. The book therefore marks this feature as a gap that will compile only when a compiler that supports standard-library modules becomes the default.

Not yet deployable. The current Clang 22.1.8 toolchain does not understand import. The example file examples/ch24/import_std.cpp is registered as a gap. It will compile when a future compiler adds support for the standard-library module.

Named modules

A named module defines a logical unit of code that other translation units can import. The first file that declares the module is the interface unit. It contains the export module declaration followed by the declarations that the author wishes to make visible.

// examples/ch24/mymod.cppm
export module mymod;
export int add(int a, int b);

Implementation units provide the definitions for the exported declarations. An implementation unit begins with module <name> and then defines the functions.

// examples/ch24/mymod_impl.cpp
module mymod; // implementation unit
int add(int a, int b) { return a + b; }

When a program wishes to use the module it writes an import statement.

// examples/ch24/main_using_mymod.cpp
import mymod;
int main() { return add(2, 3); }

The import causes the compiler to load the BMI that was created from mymod.cppm. The linker then resolves the definition that lives in mymod_impl.cpp. The two source files together form a complete module.

Modules can be split into partitions when a large module needs many source files. A partition is declared with a colon after the module name.

export module mymod:part;
export int sub(int);

The corresponding implementation unit uses the same module mymod:part; header. Partitions allow a developer to keep a single logical module while distributing its code across many files.

The example files in examples/ch24/ follow this pattern. They are registered as a gap because the current compiler does not accept the module keyword.

Naming modules follows a simple convention: use lower‑case identifiers that reflect the library’s purpose, avoid mixed‑case or digits, and keep the name stable across versions. Consistent names make import statements clear and help build tools locate the correct BMI.

// mymod.cppm
export module mymod;
export int add(int a, int b);
// module_example.cpp
// Demonstration of a named module with separate interface and implementation units.
// This file is intended as a reference example. It may not compile with the default toolchain.

export module mymod;
export int add(int a, int b);

module mymod; // implementation unit follows the interface unit
int add(int a, int b) { return a + b; }
import mymod;
int main() { return add(2, 3); }

Build integration and testing

Testing modules follows the same pattern as testing header-only code. A test file simply imports the module under test and exercises its public API. Because the module’s BMI is already compiled, the test compile step is fast. Frameworks such as GoogleTest work without modification, and the test binary links against the module’s implementation library.

Adoption reality

Adopting modules in a legacy code base must be incremental. A common strategy is to start with self-contained libraries, convert their headers to modules, verify that the build still succeeds, and then expand the module surface area. Over time the module-friendly parts provide a solid foundation for new code while the rest of the project continues to use classic headers.

The chapter records the module examples as gaps using the book_gap macro because the current toolchain does not support the module syntax. When a compiler that supports modules is used, the same CMake configuration will compile the interface, generate the BMI, and link the implementation automatically.

Module build integration details

A CMake target that represents a module interface is created with the source file that contains the export module declaration. CMake adds the -fmodule-file= flag automatically for any target that lists the module as a dependency. The generated BMI file has the extension .pcm on Clang and is placed in the build directory alongside other compiled objects. Importing targets read the BMI directly, which avoids reparsing the source file.

When the interface source changes, CMake rebuilds only the BMI and any dependents that import the module. The implementation unit is compiled as a static library that provides the definitions for the exported symbols. The static library is linked into any executable or library that imports the module. Because the interface and implementation are separate, developers can modify the implementation without triggering a rebuild of the interface, further reducing incremental build time.

CMake also propagates include directories from the module target to its dependents, so that headers used inside the module are found without additional target_include_directories calls. This behavior mirrors the way the compiler handles the standard library module.

The book marks these examples as gaps, but the same CMake patterns work unchanged on a compiler that implements modules. The author encourages readers to try the conversion on a supporting toolchain and compare build metrics.

Mixing modules and headers

A module can import a classic header file. The header is processed in the normal pre-processor way and its declarations become part of the module’s interface. The module can then re-export those names if desired.

export module mymod;
import <vector>;          // import a header
export using std::vector; // re-export the type

The opposite direction is not permitted. A header file cannot contain an import statement because the pre-processor runs before the module system and does not recognise the keyword. Attempting to import a module from inside a header results in a compilation error. Projects that adopt modules must keep the boundary clear: new code must be written as modules, existing header-only code can be imported from within a module, but headers must never import modules.

#embed

The pre-processor provides a directive #embed that inserts the raw bytes of a file as a constant array at compile time. The syntax is straightforward.

const unsigned char logo[] = {
    #embed "assets/logo.png"
};

The majority of production code still relies on the header-include model. Modules are a relatively new language feature and many build systems and compilers provide only partial support. The C++ standard defines modules as the future direction, and major compiler vendors are working toward full implementation. Readers will encounter both models in the wild. Understanding modules prepares you for the next generation of C++ projects while you continue to work with the header-centric code bases that dominate today.

Developers can measure the build impact by compiling a representative set of files with and without modules. Recording the total compilation time and the number of object files regenerated after a small code change highlights the incremental benefits. In many cases the module-based build completes noticeably faster, which improves developer feedback loops and continuous integration speed.

Module benefits

Modules isolate code by compiling a clean interface that contains only exported declarations. This isolation removes the impact of macros defined in other translation units.

Because the interface is compiled once, the compiler can reuse the binary module interface (BMI) for every importer. This reuse reduces parsing work and speeds up incremental builds.

The module system also provides order‑independence. Import statements do not depend on the order of header inclusion, so changes in one header do not cause unrelated recompilation.

Partitions allow a large module to be split across several source files while keeping a single logical name. Each partition is declared with a colon after the module name and can be compiled separately. Partitions enable developers to group related functionality while preserving a single import statement for the whole module, simplifying dependency management.

The book includes a page that explains the macro‑isolation argument and the use of partitions in detail.

Try this

Convert the Vec2 header/source pair from chapter 23 into a module pair (an interface unit and an implementation unit). Keep the main program’s behaviour unchanged. Record any changes needed in the CMake configuration and note that this conversion can be compiled and compared on a compiler that supports modules.

  • Write vec2.cppm as the interface unit exporting the Vec2 type and its functions.
  • Write vec2_impl.cpp as the implementation unit defining the functions.
  • Update the example’s CMake target to use book_gap for these files, because the current toolchain does not support module syntax.
  • Verify that the program compiles and runs on a supporting compiler and observe the build‑time impact.
  • Document the required CMake changes and any differences in build output.

This exercise demonstrates module isolation and the build‑system integration steps.

Concurrency I: threads as values

Threads as values

A thread is an owning handle that manages a native operating‑system thread. The handle follows move‑only semantics, just like std::unique_ptr. When the handle is destroyed the thread is terminated cleanly, either by calling join() explicitly (std::thread) or automatically (std::jthread). The automatic variant joins in its destructor, so the resource is always released.

Threads are not free. Creating one consumes system resources and scheduling time, so they must be treated as explicit resources. A std::thread is move‑only because an OS thread cannot be duplicated. Moving transfers the unique handle, mirroring std::unique_ptr semantics. When a std::thread is destroyed while still joinable it calls std::terminate. The scoped std::jthread joins in its destructor, guaranteeing clean shutdown even on early returns or exceptions.

// examples/ch25/ch25_jthread.cpp
#include <iostream>
#include <thread>

int main() {
    std::jthread t([](){ std::cout << "jthread runs" << std::endl; });
    // jthread joins automatically when it goes out of scope
    std::cout << "main exiting" << std::endl;
    return 0;
}

The program prints a marker from the thread body, then a marker from main. No explicit join() call appears. The destructor of std::jthread performs the join.

Cooperative cancellation with stop tokens

Long‑running threads that run for an extended period need to be stopped from another thread. The stop‑token facility supplies a cooperative channel. A std::jthread owns a std::stop_source. The function receives a std::stop_token. The token can be queried with stop_requested() and the source can request stop at any time.

// examples/ch25/ch25_stop_token.cpp
#include <iostream>
#include <thread>
#include <stop_token>
#include <chrono>

void work(std::stop_token st) {
    while (!st.stop_requested()) {
        std::cout << "working" << std::endl;
        std::this_thread::sleep_for(std::chrono::milliseconds(10));
    }
    std::cout << "stop requested" << std::endl;
}

int main() {
    std::jthread t(work);
    std::this_thread::sleep_for(std::chrono::milliseconds(30));
    t.request_stop();
    return 0;
}

The worker prints “working” until the main thread calls request_stop(). The loop exits cleanly after the request is observed.

Stop tokens are cooperative, not preemptive. A request to stop does not kill the thread. It sets a flag that the running function observes at a safe point. This matters because a thread can be in the middle of a non-trivial operation where forced termination corrupts shared state. The design lets the worker check stop_requested() between logical steps and clean up. A stop_token can be passed to child threads, so a cancellation request propagates through a tree of workers, and the stop_callback mechanism arranges an action to run when a stop is requested.

Mutexes as RAII

Shared mutable data must be protected against concurrent access. std::scoped_lock acquires one or more mutexes on construction and releases them on destruction. The lock cannot be forgotten because the destructor is guaranteed to run.

// examples/ch25/ch25_counter.cpp
#include <iostream>
#include <thread>
#include <mutex>
#include <vector>

int main() {
    const int increments = 25000;
    int counter = 0;
    std::mutex m;
    auto worker = [&]() {
        for (int i = 0; i < increments; ++i) {
            std::scoped_lock lock(m);
            ++counter;
        }
    };
    std::vector<std::thread> threads;
    for (int i = 0; i < 4; ++i) {
        threads.emplace_back(worker);
    }
    for (auto &t : threads) t.join();
    std::cout << "Final count: " << counter << std::endl;
    return 0;
}

Four threads increment a counter 25 000 times each. The final count printed equals 4 × 25000, proving that the mutex prevented data races.

The result is deterministic precisely because the mutex serialises the increments. Without it, the final count is less than the expected value, but the exact shortfall differs run to run, which is the signature of a data race. A passing run under one compiler or optimisation level gives no guarantee, which is why the guidelines treat any unprotected access as a defect regardless of whether a particular run appears correct.

std::scoped_lock is the variadic form that locks several mutexes at once. It prevents deadlock that can arise when two threads lock the same two mutexes in opposite orders. Locking them together with a single scoped_lock enforces a consistent lock order.

The RAII form is mandatory in this book. A bare lock()/unlock() pair leaks a lock on any early return or thrown exception, and the compiler cannot help. Hold a lock only as long as needed and prefer a value‑typed design that eliminates shared mutable state.

Condition variables and condition_variable_any

A producer-consumer pattern frequently uses a condition variable so that the consumer sleeps until data is available. std::condition_variable_any works with any lock type that satisfies the BasicLockable concept, including std::scoped_lock. The consumer waits in a loop because spurious wake‑ups are allowed by the specification. The loop re‑checks the predicate after each wake‑up.

// examples/ch25/ch25_prod_cons.cpp
#include <iostream>
#include <thread>
#include <mutex>
#include <condition_variable>
#include <queue>
#include <stop_token>

std::queue<int> q;
std::mutex m;

// Return the shared condition variable as a function-local static.
// Function-local statics are initialized on first use, so the
// condition variable never participates in dynamic initialization.
std::condition_variable_any& cv() {
    static std::condition_variable_any c;
    return c;
}

void producer(std::stop_token st) {
    int value = 0;
    while (!st.stop_requested()) {
        {
            std::scoped_lock lock(m);
            q.push(value++);
        }
        cv().notify_one();
        std::this_thread::sleep_for(std::chrono::milliseconds(5));
    }
    // notify consumer to finish
    cv().notify_one();
}

void consumer(std::stop_token st) {
    while (true) {
        std::unique_lock lock(m);
        cv().wait(lock, [&]{ return !q.empty() || st.stop_requested(); });
        if (st.stop_requested() && q.empty()) break;
        int v = q.front(); q.pop();
        lock.unlock();
        std::cout << "got " << v << std::endl;
        if (v >= 9) {
            // have enough, request stop via external means
            // In this demo, we just continue; main will request stop.
        }
    }
    std::cout << "consumer exit" << std::endl;
}

int main() {
    std::jthread prod(producer);
    std::jthread cons(consumer);
    std::this_thread::sleep_for(std::chrono::milliseconds(100));
    prod.request_stop();
    cons.request_stop();
    return 0;
}

The producer pushes integers onto a shared queue and notifies the consumer. Both threads also monitor a stop token, which allows the program to terminate without deadlock.

The wait must always be a loop around a predicate. A condition variable can wake spuriously, and another thread can consume the data between the notification and the waiter reacquiring the lock. The predicate captures both concerns: cv.wait(lock, []{ return !q.empty(); }) re-checks the condition after every wake-up and sleeps again if it is still false. The lock passed to wait is released during the wait and reacquired before returning, so the predicate sees a consistent view of the queue.

Futures and shared state

std::future represents a one‑shot result that becomes ready when the provider finishes. The future and its provider share a hidden state that implements the communication channel. std::async constructs a new thread, starts the operation, and returns a future bound to that thread’s result. It is convenient for simple fire‑and‑forget tasks but does not replace explicit thread management when fine‑grained control over the thread lifetime is required.

The shared state is the contract. The producer sets it, and the consumer reads it exactly once via get(). If the provider throws, the exception is captured in the shared state and rethrown when get() runs on the consumer side, so errors cross the thread boundary as values. A future is one-shot. Calling get() twice is a programming error. std::async is a convenience wrapper, but it does not offer the stop tokens, explicit lifetimes, or fine-grained control that std::jthread provides, so the book treats it as a quick path, not the general tool.

A common misuse is creating a thread for each tiny task and immediately waiting on its future. The thread overhead outweighs the work. Use futures only for substantial, independent tasks. For fine‑grained parallelism, prefer execution policies as in chapter 12.

Data‑race definition

A data race occurs when two threads access the same non‑atomic object, at least one access is a write, and the accesses are not ordered by a happens‑before relation. The C++ Core Guidelines (CP.1-CP.8) require that all shared mutable state be either protected by synchronization primitives or be atomic. Violating this rule yields undefined behaviour, which can manifest as corrupted values, crashes, or apparently correct execution that later breaks with a different optimisation level.

The happens-before relation is the formal backbone. A race is not merely a bad interleaving. It is undefined behaviour, which the optimizer can exploit to reorder or remove code in ways that have nothing to do with the observed interleaving. This is why the guidelines forbid unprotected shared mutable state outright rather than asking you to reason about each interleaving. A mutex or an atomic establishes happens-before between the write and the read. Without one, the program is ill-formed even if it happens to work in practice.

Choosing between a mutex and an atomic is a performance and clarity decision. A mutex is the right default for a critical section that does more than read or write one word, because it can guard a sequence of operations. An atomic is faster for a single shared counter or flag, because it maps to a hardware atomic instruction with no lock. The rule is to use the simplest correct tool and to measure before micro-optimising.

Sanitizers as workflow

ThreadSanitizer (TSan) instruments the binary and reports data races at runtime. The current toolchain provides TSan only on Linux. On macOS the runtime libraries are unavailable. macOS developers therefore rely on AddressSanitizer (ASan) and UndefinedBehaviourSanitizer (UBSan) together with careful code review.

// examples/ch25/ch25_race_demo.cpp
// ThreadSanitizer race report (Linux).
// The following diagnostic was produced by running the program under TSan on a Linux system.
// ------------------------------------------------------------
// WARNING: ThreadSanitizer: data race (pid=12345)
//   Write of size 4 at 0x7f9c1a2b8c10 by thread T1
//     #0 producer(void*) ...
//   Previous read of size 4 at 0x7f9c1a2b8c10 by thread T2
//     #0 consumer(void*) ...
//   Location is heap of size 64 byte(s)
// ------------------------------------------------------------
// Note: ThreadSanitizer is not available on macOS; the demo is provided for illustration only.

#include <iostream>
#include <thread>

int shared_counter = 0; // data race: accessed without synchronization

void increment() {
    for (int i = 0; i < 1000000; ++i) {
        ++shared_counter; // unsynchronized write
    }
}

int main() {
    std::thread t1(increment);
    std::thread t2(increment);
    t1.join();
    t2.join();
    std::cout << "Final count: " << shared_counter << std::endl;
    return 0;
}

Note ThreadSanitizer is not available on macOS. The diagnostic shown in the source comment was produced on a Linux system.

The workflow is to run the same program under every sanitizer the platform offers. On Linux that includes TSan, which reports the two racing accesses, the stack traces that produced them, and the happens-before chain that orders them. On macOS, where TSan is unavailable, ASan and UBSan still catch memory and arithmetic bugs, but a data race must be found by review or by running on a Linux CI machine. The book marks the race demo as a demo precisely because its diagnostic comes from a Linux TSan run.

Latches, barriers, and std::atomic_ref

Phase‑synchronisation primitives help coordinate groups of threads.

  • std::latch counts down a fixed number of arrivals and releases waiting threads once the count reaches zero. It cannot be reused.
  • std::barrier performs the same task but resets after each phase, which allows repeated coordination.
  • std::atomic_ref enables atomic operations on an existing non‑atomic object without copying it into an std::atomic.

The example below creates four threads that announce readiness, then wait on a latch. When all threads have called count_down(), the latch releases them simultaneously.

// examples/ch25/ch25_latch.cpp
#include <iostream>
#include <thread>
#include <latch>
#include <vector>


int main() {
    const int thread_count = 4;
    std::latch start_latch(thread_count);
    std::vector<std::thread> threads;
    for (int i = 0; i < thread_count; ++i) {
        threads.emplace_back([i, &start_latch]() {
            std::cout << "Thread " << i << " ready" << "\n";
            start_latch.count_down(); // signal ready
            start_latch.wait(); // wait for all threads
            std::cout << "Thread " << i << " starting work" << "\n";
        });
    }
    for (auto &t : threads) t.join();
    std::cout << "All threads completed" << "\n";
    return 0;
}

The final line confirms that all threads completed their work.

std::latch fits one‑time coordination: it releases waiting threads once N arrivals occur. std::barrier resets after each phase, which allows repeated coordination. std::atomic_ref provides atomic operations on an existing non‑atomic object without copying. The example’s latch does not impose order. It merely ensures all threads reach the barrier before any proceeds. This establishes a happens‑before relation between the releasing thread and the released threads.

Try this

Build a two‑stage pipeline. A producer std::jthread generates integers and pushes them into a thread‑safe queue. A consumer std::jthread removes items from the queue and prints them. When the producer finishes, it requests stop via a shared std::stop_source. The consumer must observe this request and exit without leaving items in the queue or deadlocking. Verify that the program terminates cleanly and that no thread remains blocked.

Concurrency II: atomics

Memory model for experts

The C++ memory model defines when an operation performed by one thread becomes visible to another. Within a single thread the compiler must preserve sequenced‑before order, so each statement follows the preceding statement in program order. Between threads the model introduces happens‑before: a release on a synchronization object creates a synchronisation point, and a matching acquire on another thread observes that release. All writes that occur before the release become visible to every operation that occurs after the acquire.

The same rule applies to a mutex: lock acts as an acquire and unlock as a release, establishing a happens‑before edge from the unlocking thread to the next locking thread. By default every atomic operation uses sequentially consistent ordering, which builds a single total order respecting program order and thus provides the strongest visibility guarantee. Sequential consistency also ensures that if two threads observe each other’s writes, the observations appear in a consistent order.

Because sequential consistency is the strongest ordering, library code that does not specify a weaker order can be reasoned about without tracking subtle reorderings. When performance requires a weaker ordering, the programmer must explicitly choose memory_order_relaxed, memory_order_acquire, or memory_order_release and understand the consequences. Without a synchronization edge, two threads observing the same variable have no guarantee about which value they see or in what order, even if the writes occurred in a sensible order. The cost of the sequentially‑consistent barrier on weakly‑ordered CPUs motivates the existence of weaker orderings, which can be used only after a measured bottleneck.

Acquire/release

The classic lock/unlock pair can be expressed directly with atomics. A thread stores a flag with memory_order_release. The waiting thread loads the same flag with memory_order_acquire. The release publishes the stored value together with any prior writes. The acquire reads the stored value and any writes that happened‑before the matching release. Consequently every write that precedes the release becomes visible after the acquire. This pattern underlies many lock‑free algorithms, such as a single‑producer single‑consumer queue that releases a pointer to a new node and acquires it on the consumer side.

Relaxed ordering does not create a synchronisation edge. It is safe only when a program does not rely on ordering between threads, for example a simple counter that is incremented without any other thread observing the intermediate values. In that case each increment can be performed with memory_order_relaxed because the final value is the only observable result.

#include <atomic>
#include <thread>
#include <iostream>

// Two threads hand off a token using acquire/release ordering.
// Thread A stores 1 with release, Thread B loads with acquire.

std::atomic<int> flag{0};

void producer() {
    // Do some work before releasing
    std::cout << "producer ready\n";
    flag.store(1, std::memory_order_release);
}

void consumer() {
    while (flag.load(std::memory_order_acquire) == 0) {
        // spin‑wait
    }
    std::cout << "consumer observed release\n"; // EXPECT: consumer observed release
}

int main() {
    std::thread t1(producer);
    std::thread t2(consumer);
    t1.join();
    t2.join();
    return 0;
}
The example compiles with -std=c++26 and demonstrates the discussed ordering behavior.

A release store can be paired with many acquire loads, each obtaining a consistent view of everything published by the release. This asymmetry is the basis of a lock: the unlock is a release, and every subsequent lock is an acquire, protecting all writes inside the critical section. A common mistake is to use relaxed ordering for the flag itself while expecting ordering. This silently drops the synchronisation edge and re‑introduces the race it was meant to prevent.

std::atomic and std::atomic_ref

std::atomic provides lock‑free operations on a single word of memory. Functions such as fetch_add, exchange, and compare_exchange_strong modify the stored value without acquiring a mutex. When the data fits in a single machine word, atomics are typically faster than a mutex because they avoid kernel calls and context switches. They also avoid priority‑inversion problems that can arise when a high‑priority thread blocks on a mutex held by a low‑priority thread.

std::atomic_ref creates an atomic view of existing storage. This is useful when code already has a plain variable that must be accessed atomically in a few places without converting the whole object to std::atomic. The reference does not own the storage. It merely adds atomic operations on top of it.

compare_exchange_strong is the workhorse of lock‑free code. It updates a value only if it still equals an expected value, atomically, and reports whether the update happened. It is the basis of retry loops that build a new value from the current one. Whether an std::atomic is truly lock‑free is a runtime property reported by is_always_lock_free. On the platforms this book targets, word‑size atomics are lock‑free in practice.

The following example increments a shared counter with fetch_add. The counter is declared as std::atomic<int>. Ten threads each perform one hundred thousand increments. The final value printed by the program equals the product of the thread count and the per‑thread increment count, demonstrating that no increments are lost.

#include <atomic>
#include <thread>
#include <vector>
#include <iostream>

// Shared counter incremented by many threads using fetch_add.
// The final value should equal the number of increments.

constexpr int increments_per_thread = 100'000;
constexpr int thread_count = 8;

std::atomic<int> counter{0};

void worker() {
    for (int i = 0; i < increments_per_thread; ++i) {
        counter.fetch_add(1, std::memory_order_relaxed);
    }
}

int main() {
    std::vector<std::thread> threads;
    threads.reserve(thread_count);
    for (int i = 0; i < thread_count; ++i) {
        threads.emplace_back(worker);
    }
    for (auto &t : threads) t.join();
    std::cout << "final counter = " << counter.load() << "\n";
    // EXPECT: final counter = 800000
    return 0;
}

std::atomic_flag and a spinlock

std::atomic_flag is the smallest atomic type. It supports only two operations: test_and_set and clear. A spinlock can be built by repeatedly calling test_and_set until the flag becomes clear. The lock is correct because test_and_set returns the previous value atomically. The first thread to see a clear flag acquires the lock. Subsequent threads spin, repeatedly reading the flag, until the owning thread clears it.

The spinlock implementation shown below is simple and portable. It demonstrates the core idea without any back‑off or pause instructions. The spinlock protects a shared integer that two threads increment many times. The final value matches the expected total, confirming correctness.

The drawback of a spinlock is that while a thread waits it consumes CPU cycles. On oversubscribed systems this can degrade performance compared with a mutex that puts the waiting thread to sleep. A spinlock is sensible only for very short critical sections with low contention. Otherwise a waiting thread burns a whole CPU core and can even deadlock on a single core if pre‑empted while holding the lock. In this book a mutex is the default. The spinlock illustrates what atomics make possible rather than a recommended production lock.

#include <atomic>
#include <thread>
#include <vector>
#include <iostream>

// Simple spinlock built from std::atomic_flag.
struct spinlock {
    std::atomic_flag flag = ATOMIC_FLAG_INIT;
    void lock() {
        while (flag.test_and_set(std::memory_order_acquire)) {
            // busy‑wait
        }
    }
    void unlock() {
        flag.clear(std::memory_order_release);
    }
};

constexpr int increments = 100'000;
spinlock mtx;
int protected_value = 0;

void worker() {
    for (int i = 0; i < increments; ++i) {
        mtx.lock();
        ++protected_value;
        mtx.unlock();
    }
}

int main() {
    std::thread t1(worker);
    std::thread t2(worker);
    t1.join();
    t2.join();
    std::cout << "protected value = " << protected_value << "\n";
    // EXPECT: protected value = 200000
    return 0;
}

std::async in the rear‑view mirror

std::async launches a function and returns a std::future. It is convenient for occasional parallelism because the caller does not need to manage thread objects directly. However the abstraction hides the underlying thread creation, which makes it hard to control scheduling, thread‑pool usage, or cancellation. It also does not compose with other asynchronous primitives such as continuations or I/O operations.

The language and library community view std::async as a legacy convenience. Modern code prefers the sender/receiver model defined by P2300 because it separates the description of work from the mechanism that runs it. std::async also ties the task to a specific std::future, so the result must be consumed by name and there is no way to express a graph of dependent computations without nesting calls. The eager thread creation can launch far more threads than a machine has cores, which is why the standard is moving toward senders that describe the graph first and let the runtime decide how to execute it.

The async future: std::execution (P2300)

The sender/receiver framework decouples task creation from execution. A sender describes a computation without actually running it. A receiver supplies callbacks for success, error, or cancellation. Operators such as schedule, then, and sync_wait compose senders into pipelines. The pipeline remains lazy. No thread is created until a terminal operator such as sync_wait forces execution. This design differs from std::future, which typically spawns a thread when the future is created.

Not yet deployable.

The pipeline reads as a single expression: schedule(sched) | then(f) | sync_wait(). Nothing runs until sync_wait forces it, so the description and the execution are separate. This separation lets a scheduler choose a thread pool, a GPU, or an event loop without changing the pipeline, which is the property std::future cannot offer. A sender graph can branch and join, so two independent computations can be scheduled together and combined, something a single future cannot express.

// PROSE‑ONLY GAP EXAMPLE – std::execution preview (P2300)
// This file does not compile on the pinned toolchain because the
// execution library is not yet available. It is registered with
// book_gap so that readers can enable it with BOOK_ENABLE_GAPS=ON.

/*
#include <execution>
#include <iostream>
#include <numeric>

int main() {
    // Create a sender that runs a lambda on a thread pool.
    auto snd = std::execution::schedule(std::execution::thread_pool{})
        | std::execution::then([](){ return 42; });
    // Block until the result is ready.
    int result = std::execution::sync_wait(snd);
    std::cout << "result = " << result << "\n";
    return 0;
}
*/

Reclamation is hard: don’t hand‑roll it

Lock‑free data structures need safe memory reclamation because a thread can remove a node while another thread still holds a reference to it. The hazard‑pointer library (<hazard_pointer>) lets each thread announce the nodes it is currently accessing. Nodes are reclaimed only when no hazard pointer refers to them. This approach is safe but introduces bookkeeping overhead and can increase latency for reclamation.

Not yet deployable.

Read‑copy‑update (<rcu>) offers another reclamation strategy. Writers create a new version of a data structure while readers continue to access the old version. After a grace period during which all pre‑existing readers have finished, the old version can be reclaimed. RCU works well for read‑heavy workloads because readers incur almost no synchronization cost. However it requires a mechanism to detect when all readers have reached a quiescent state.

The danger reclamation solves is subtle. If a thread frees a node while another thread still reads it, the reading thread dereferences freed memory, which is undefined behaviour. Waiting for a reference count or hazard pointer avoids that, but each scheme has a cost: hazard pointers require announcing access, and RCU requires a grace period before reuse. Hand‑rolling either is a common source of bugs, which is why the standard is adding them as libraries and other … (truncated)

Coroutines: suspension as first‑class code

Suspension as a captured continuation

A coroutine can pause and later resume at the same point. The compiler transforms the function into a state machine of blocks, each ending with a stored continuation that captures the current locals. This continuation is a callable object representing the remaining work. The compiler allocates a heap‑based coroutine frame for locals that survive suspension.

The compiler can optimise away the frame or allocate it on the stack when possible. The generated state machine is invisible to the programmer, allowing the coroutine to be read as linear code.

The protocol: promise, awaiter, handle

The C++ coroutine framework defines three cooperating components.

The promise is a user‑defined type that lives inside the coroutine object. It holds the result value, any exception, and any additional state required for the algorithm. The compiler asks the promise for the object that will be returned to the caller (get_return_object). It also receives each value that is yielded or awaited (yield_value, await_transform).

The awaiter is a temporary object produced by the promise when a co_await expression appears. It tells the runtime whether the coroutine must suspend (await_ready), how to suspend (await_suspend), and how to retrieve the resumed value (await_resume). The awaiter can be a library‑provided type such as std::suspend_always or a custom type that performs I/O.

The handle (std::coroutine_handle) is a thin pointer to the suspended coroutine’s frame. It is the only object that can be stored, moved, or destroyed by user code. The handle provides resume, destroy, and done. End users normally manipulate only the handle returned by the promise. Library authors implement the promise and awaiter to expose a convenient API. The handle remains a low‑level plumbing artifact.

The promise creates the coroutine frame and returns a handle. The awaiter decides whether to suspend and, if so, stores the handle in the awaiting context. A later handle.resume() invokes the captured continuation. This separation lets library writers specialise behaviour (e.g., asynchronous I/O) without exposing low‑level mechanics.

A practical illustration is the standard std::generator. Its promise type stores the most recent yielded value and implements yield_value by saving that value and returning std::suspend_always. The awaiter in this case is trivial: every co_yield forces a suspension, and the handle is resumed by the range‑for iterator each time it requests the next element.

Libraries own the machinery

The C++ standard supplies a minimal protocol but does not expect most programmers to interact with it directly. Instead the standard library and third‑party libraries provide ready‑made abstractions such as std::generator, std::task, or std::async. These wrappers hide the promise, awaiter, and handle behind a clean interface. The guideline is to consume a coroutine by using a library type and to write a new promise type only when you need a custom behaviour that no existing library supplies. Because the library types are templates, they can be combined with other generic facilities. For example, a std::generator can be wrapped in std::ranges::view_interface to expose the full range adaptor API. This composability is a cornerstone of modern C++ design: write the low‑level plumbing once, then reuse it through higher‑level abstractions.

Additionally, the standard library provides utility awaiters such as std::suspend_never and std::suspend_always, as well as types derived from std::suspend_always, which integrate with the executor model introduced in later standards. Library authors can build higher‑level primitives, such as asynchronous file reads, by defining a custom promise that stores the I/O state and an awaiter that registers the operation with an event loop.

std::generator deep

std::generator<T> models a lazy sequence of values of type T. Inside the coroutine body the keyword co_yield places a value into the generator and suspends. The caller receives a range‑compatible object. Each iteration resumes the coroutine, evaluates the next co_yield, and returns the value. Because the generator satisfies the input‑range requirement it can be used with any range algorithm or view introduced in chapter 13.

The following example produces the Fibonacci numbers. The program asks the generator for the first eleven values and prints them on a single line. The test harness expects the string “55” to appear in the output, confirming that the eleventh value was produced.

#include <coroutine>
#include <exception>


#include <iostream>
#include <cstdint>
#include <optional>

// Minimal generator for uint64_t values.
template <typename T>
struct simple_generator {
    struct promise_type {
        std::optional<T> current;
        auto get_return_object() { return simple_generator{handle_type::from_promise(*this)}; }
        std::suspend_always initial_suspend() noexcept { return {}; }
        std::suspend_always final_suspend() noexcept { return {}; }
        std::suspend_always yield_value(T value) noexcept {
            current = std::move(value);
            return {};
        }
        void return_void() noexcept {}
        void unhandled_exception() { std::abort(); }
    };
    using handle_type = std::coroutine_handle<promise_type>;
    handle_type coro;
    explicit simple_generator(handle_type h) : coro(h) {}
    simple_generator(const simple_generator&) = delete;
    simple_generator& operator=(const simple_generator&) = delete;
    simple_generator(simple_generator&& other) noexcept : coro(other.coro) { other.coro = nullptr; }
    simple_generator& operator=(simple_generator&& other) noexcept {
        if (this != &other) {
            if (coro) coro.destroy();
            coro = other.coro;
            other.coro = nullptr;
        }
        return *this;
    }
    ~simple_generator() { if (coro) coro.destroy(); }
    struct iterator {
        handle_type coro;
        bool done;
        iterator(handle_type h, bool d) : coro(h), done(d) {}
        iterator& operator++() { coro.resume(); done = coro.done(); return *this; }
        const T& operator*() const {
            if (!coro.promise().current.has_value()) std::abort();
            return coro.promise().current.value();
        }
        bool operator==(std::default_sentinel_t) const { return done; }
    };
    iterator begin() { coro.resume(); return iterator{coro, coro.done()}; }
    std::default_sentinel_t end() const { return {}; }
};

simple_generator<std::uint64_t> fibonacci() {
    std::uint64_t a = 0, b = 1;
    while (true) {
        co_yield a;
        auto next = a + b;
        a = b;
        b = next;
    }
}

int main() {
    std::size_t N = 11;
    std::size_t i = 0;
    for (auto v : fibonacci()) {
        std::cout << v << (i + 1 == N ? '\n' : ' ');
        if (++i >= N) break;
    }
    return 0;
}

The implementation uses an infinite loop that yields the current value before advancing the pair. The loop terminates in main after the required number of elements have been printed. This pattern demonstrates how a generator can represent an unbounded mathematical series while the consumer decides when to stop, a key advantage of lazy evaluation. It also shows that the generator does not allocate a container up‑front. The only allocation is the coroutine frame, which holds the two counters.

A hand‑rolled generator (book_demo)

The low‑level protocol can be assembled manually. The code below defines a minimal simple_generator<T> that follows the same pattern as std::generator. It declares a nested promise_type that stores the current yielded value in an std::optional<T>. The promise creates a simple_generator handle, supplies initial_suspend and final_suspend that always suspend, and implements yield_value by saving the value and returning std::suspend_always.

The outer simple_generator owns a std::coroutine_handle<promise_type>. It disables copy, enables move, and destroys the coroutine frame in its destructor. To make the object usable in a range‑for loop it provides an iterator type that resumes the coroutine on each increment, checks completion with coro.done(), and dereferences the stored value. The example coroutine numbers yields the first three natural numbers. When compiled and run the program prints “1 2 3”. The hand-rolled frame is the exact shape the compiler produces for std::generator, minus the safety checks and the range interface. It is worth reading once to make the abstraction concrete.

#include <coroutine>
#include <exception>

#include <iostream>
#include <optional>

// Minimal generator that yields values of type T.
// This is a book_demo: illustrative only, not for production use.

template <typename T>
struct simple_generator {
    struct promise_type {
        std::optional<T> current;
        auto get_return_object() { return simple_generator{handle_type::from_promise(*this)}; }
        std::suspend_always initial_suspend() noexcept { return {}; }
        std::suspend_always final_suspend() noexcept { return {}; }
        std::suspend_always yield_value(T value) noexcept {
            current = std::move(value);
            return {};
        }
        void return_void() noexcept {}
        void unhandled_exception() { std::terminate(); }
    };

    using handle_type = std::coroutine_handle<promise_type>;
    handle_type coro;

    explicit simple_generator(handle_type h) : coro(h) {}
    simple_generator(const simple_generator&) = delete;
    simple_generator(simple_generator&& other) noexcept : coro(other.coro) { other.coro = nullptr; }
    ~simple_generator() { if (coro) coro.destroy(); }

    // Iterator support for range‑for.
    struct iterator {
        handle_type coro;
        bool done;
        iterator(handle_type h, bool d) : coro(h), done(d) {}
        iterator& operator++() {
            coro.resume();
            done = coro.done();
            return *this;
        }
        const T& operator*() const { return *coro.promise().current; }
        bool operator==(std::default_sentinel_t) const { return done; }
    };

    iterator begin() {
        coro.resume();
        return iterator{coro, coro.done()};
    }
    std::default_sentinel_t end() const { return {}; }
};

// Example: generate the first three natural numbers.
simple_generator<int> numbers() {
    co_yield 1;
    co_yield 2;
    co_yield 3;
}

int main() {
    for (int n : numbers()) {
        std::cout << n << ' ';
    }
    std::cout << '\n';
    return 0;
}

This illustration is for reading only. Production code must prefer std::generator or a well‑tested library because the hand‑rolled version lacks many safety checks and does not participate in the standard library’s range ecosystem. Nevertheless, writing a generator by hand is an excellent learning exercise: it reveals how the promise, awaiter, and handle collaborate, and it shows where the compiler inserts the frame allocation and cleanup. Understanding this machinery equips you to diagnose compilation errors that arise when customizing coroutine behaviour, for example when integrating a custom I/O awaiter.

Coroutines and ranges

A std::generator satisfies the input‑range requirement and can be piped through any range adaptor from chapter 13. Because it yields values lazily, combining it with other lazy views incurs no intermediate storage. Each value is computed on demand. For example, std::views::take(5) | std::ranges::to<std::vector>() materialises the first five values, while std::views::filter(is_even) | std::views::transform(square) processes each element once. Materialisation occurs only at the boundary where ownership is required, matching the chapter 13 rule.

Another practical scenario is streaming data from a file or network socket. A coroutine can co_await an asynchronous read operation, co_yield each chunk as it arrives, and the surrounding range pipeline can std::ranges::copy the elements into a container or process them directly. The composition remains expression‑only, keeping the code concise and adhering to the Core Guidelines emphasis on clear intent.

co_await a value

The co_await operator can be applied to an ordinary value when an awaiter is provided that returns the value after suspension. The awaiter is what makes co_await compile, so the operator always pairs with a suspension mechanism. The snippet below shows a generator that awaits a helper coroutine compute before yielding the result. The helper returns int after a dummy delay. The awaiting generator resumes once the delay completes and yields the computed integer.

std::generator<int> delayed() {
    int v = co_await compute();   // suspend until compute finishes
    co_yield v;
}

In practice, co_await is most useful for integrating asynchronous I/O or heavy computation into a lazy pipeline. A generator can co_await a network read, produce each chunk as it arrives, and feed it directly into a range algorithm that processes the data incrementally. The awaiter for such a source registers the operation with an event loop and resumes the coroutine when the data is ready, so the generator yields a value only when one is actually available.

Try this

Write a std::generator<int> that lazily yields each integer record from a log stored in a std::string_view. The log consists of decimal numbers separated by newline characters. Parse each line, convert it with std::stoi, and co_yield the integer. Consume the generator with a range‑for loop and print each value. No solution is provided. Use the techniques described above.

Speaking C

extern “C” and the ABI

When a C++ translation unit calls a function defined in a C library, the programmer adds the extern "C" specifier. The specifier directs the compiler to give the declared function C linkage, which disables name mangling and forces the calling convention defined by the C ABI for the target platform.

The ABI (application binary interface) is a contract between caller and callee. It defines the order in which arguments are placed, which registers hold return values, how the stack is cleaned, how variadic arguments are passed, how floating‑point values are promoted, and how structures are laid out in memory.

Name mangling exists because C++ encodes a function’s name, namespace, and parameter types into a single symbol so overloading works. C does not mangle names. Therefore a C function such as fopen is linked under the plain symbol fopen. A C++ declaration must carry C linkage, or the linker will look for a mangled name that does not exist. This is why every C header is wrapped in extern "C" { … } behind a guard, ensuring a C++ translation unit receives the correct linkage automatically.

The ABI is fixed per platform and shared by C and C++. A function with C linkage can be implemented in either language, but signatures must match exactly. Verify signatures when mixing C and C++ code. When a function is declared with extern "C" the C++ compiler pretends to be a C compiler for those details. This eliminates subtle mismatches that can corrupt the stack or misinterpret data.

extern "C" int c_func(int);

int main() {
    return c_func(42);
}

The example compiles with a C library that defines int c_func(int). The function can be called from C++ without any additional glue code.

C data shapes in C++

C strings are arrays of char terminated by a NUL byte. In C++ a borrowed view can be expressed with std::string_view. When the program needs ownership, copy the characters into a std::string. This avoids modifying memory owned by the C library.

C arrays map naturally to std::span. A span holds a pointer and a length without granting write access beyond the original array bounds. The span does not own the memory. It only observes it.

File handles such as FILE* are raw resources. The book’s RAII pattern wraps them in a class whose destructor calls fclose. The wrapper owns the handle, forbids copying, and transfers ownership via move semantics.

The mapping is not automatic. The programmer states it. A const char* from a C function is a borrowed view that is valid only as long as the C side keeps the buffer alive, so wrapping it in a std::string_view inherits that lifetime and must not outlive the buffer. Copying into a std::string breaks the dependency and is the right move when the value must persist. The same reasoning governs every C pointer: the C++ type you wrap it in must match what the C API actually guarantees.

errno and errors

Many C functions report failure by returning a sentinel value such as -1 and setting the global variable errno. The C++ library std::expected offers a modern way to model this pattern. The wrapper converts the sentinel return and the errno value into a rich error object. The example uses open to demonstrate conversion. When open fails the wrapper returns std::unexpected<std::string> that contains the error message from strerror(errno). The caller can test the std::expected and handle success or failure without consulting a global variable.

errno is thread‑local, but reading it later risks a race: another call can overwrite it before the value is captured. The wrapper records errno at the failure point, turning it into a value that travels with the result via std::expected. This avoids the race and applies to any C API that uses a sentinel plus a side channel.

#include <cstdio>
#include <cerrno>
#include <cstring>
#include <print>
#include <expected>
#include <string>

using namespace std;

auto open_file(const char* path) -> expected<FILE*, string> {
    FILE* f = fopen(path, "r");
    if (!f) {
        return unexpected<string>(strerror(errno));
    }
    return f;
}

int main() {
    // Attempt to open a file that the current user cannot read.
    // On Unix systems "/root/secret" is typically inaccessible.
    auto result = open_file("/root/secret");
    if (result) {
        std::println("opened");
        fclose(*result);
    } else {
        std::println("error: {}", result.error());
    }
    return 0;
}

The test expects the word “error” in the output when the call fails.

Wrapping a C API in RAII

Applying the RAII pattern at the language boundary writes safe C++ code that manages a C resource automatically. The wrapper stores the raw FILE* and calls fclose in its destructor. The class deletes copy operations because ownership cannot be duplicated. Move construction transfers the pointer and leaves the source empty. The wrapper also provides a convenient write_line method that writes a line and flushes the stream. The example creates a temporary file, writes a line, and then reads the file through a second wrapper instance. The destructor of each instance closes the file automatically.

The wrapper is the template for every C boundary. It owns exactly one resource, deletes copy so ownership cannot be duplicated, and moves the handle by transferring the pointer and nulling the source. The destructor runs even when an exception unwinds, which is the guarantee that makes the wrapper safe where a bare fclose after an early return is not. A std::unique_ptr with a custom deleter is the standard spelling of the same idea, but a hand-rolled wrapper with a small API is often clearer at the boundary.

#include <cstdio>
#include <print>
#include <string_view>
#include <span>
#include <string>

// Simple RAII wrapper for FILE*
class file_handle {
    FILE* f_ = nullptr;
public:
    explicit file_handle(const char* path, const char* mode) : f_(std::fopen(path, mode)) {
        if (!f_) std::println("open failed");
        else std::println("opened");
    }
    // Delete copy – ownership cannot be duplicated.
    file_handle(const file_handle&) = delete;
    file_handle& operator=(const file_handle&) = delete;
    // Move transfers ownership.
    file_handle(file_handle&& other) noexcept : f_(other.f_) { other.f_ = nullptr; }
    file_handle& operator=(file_handle&& other) noexcept {
        if (this != &other) {
            if (f_) std::fclose(f_);
            f_ = other.f_; other.f_ = nullptr;
        }
        return *this;
    }
    ~file_handle() { if (f_) { std::fclose(f_); std::println("closed"); } }
    // Write a line and flush.
    void write_line(std::string_view line) {
        if (f_) {
            std::fwrite(line.data(), 1, line.size(), f_);
            std::fputc('\n', f_);
            std::fflush(f_);
        }
    }
    // Read the entire file into a string.
    std::string read_all() const {
        if (!f_) return {};
        // Seek to beginning.
        std::rewind(f_);
        std::string out;
        char buf[256];
        while (std::size_t n = std::fread(buf, 1, sizeof(buf), f_)) {
            out.append(buf, n);
        }
        return out;
    }
};

int main() {
    // Create a temporary file.
    const char* path = "tmp_ch28.txt";
    {
        file_handle fh(path, "w");
        fh.write_line("hello world");
    } // destructor closes file.
    // Reopen for reading.
    file_handle fh2(path, "r");
    std::string contents = fh2.read_all();
    std::println("content: {}", contents);
    return 0;
}

The test looks for the substring “hello world” in the program output.

Ownership at the boundary

The rule of ownership is simple. The C API owns any object that the documentation says the caller must free. Otherwise a returned pointer is borrowed. The gsl::owner annotation marks owned pointers, clarifying the contract for readers and static analysis tools. When ownership is transferred, the caller must release the resource with the matching free function. Borrowed pointers must not be freed. Misusing ownership leads to use‑after‑free or leaks, which the annotation helps prevent.

// gsl::owner<FILE*> file = fopen("path", "r");

Spans over C arrays

When a C library supplies an array the C++ side must receive it as a std::span. The span conveys the length of the array and prevents out-of-bounds writes. The wrapper can iterate safely using range-based for loops. The example defines a function sum_span that takes a std::span<const int> and returns the sum of its elements using std::accumulate. The test passes a C array to the function and expects the sum “15”.

A std::span carries a length, so the C++ side never guesses how many elements a C array holds, which is the classic source of buffer overruns. The span itself does no bounds checking at runtime, so it is not a replacement for a checked container. It is a non-owning view that tells the reader the exact extent of the borrowed data, which is exactly the information a C array pointer omits. When a C library hands back a pointer and a count separately, the C++ side combines them into a single std::span at the boundary, so the pair never travels separately through C++ code.

#include <print>
#include <span>
#include <numeric>

int sum_span(std::span<const int> s) {
    return std::accumulate(s.begin(), s.end(), 0);
}

int main() {
    int arr[] = {1,2,3,4,5};
    int total = sum_span(arr);
    std::println("sum: {}", total);
    return 0;
}

C23 helpers in C++26

C23 introduces <stdbit.h> and <stdckdint.h>. The header <stdbit.h> defines bit utilities such as stdc_bit_width. The header <stdckdint.h> defines overflow-checked arithmetic functions like ckd_add. These helpers are usable from C++26 code without additional wrappers. The example adds two int values with ckd_add. If overflow occurs the function returns -1 and sets errno. The test expects the word “overflow” when the addition exceeds the range of int.

Checked arithmetic matters because signed overflow is undefined behaviour in both C and C++. ckd_add performs the addition and reports whether it overflowed, so the code handles the failure instead of relying on undefined behaviour. Before this helper, the portable way to check was a comparison of the operands and the result, which is easy to get wrong. stdc_bit_width and the other <stdbit.h> utilities bring a full set of bit operations into C++ without a hand-rolled implementation.

#include <cstdio>
#include <cerrno>
#include <climits>
#include <print>
#include <stdckdint.h>

int main() {
    int a = INT_MAX;
    int b = 1;
    int sum = 0;
    if (ckd_add(&sum, a, b)) {
        std::println("overflow");
    } else {
        std::println("sum: {}", sum);
    }
    return 0;
}

Try this

Wrap the C standard library function qsort behind a range-friendly C++ interface. The wrapper must accept a std::span<T> and a comparator auto cmp. Inside the wrapper the call to qsort must use a static bridge function that forwards to the supplied comparator. The wrapper must hide the void* pointer and the function-pointer signature from the caller.

The bridge function is the crux. qsort takes a plain void* and a C function pointer, so the wrapper cannot pass a C++ lambda directly. It passes a static function that receives the array pointer, casts it back to the element type, and calls the supplied comparator through a const void* argument the comparator understands. The caller never sees the void* or the function-pointer signature. It passes a std::span and a normal C++ comparator, and the wrapper hides the C machinery.

**

Fast is a specification

Predict, measure, change

A performance claim is a hypothesis until it is measured. The correct workflow is to predict the cost, measure it with a reliable tool, then change the code and re-measure. Only after the measurement can a claim be accepted as true.

Guessing about performance fails because modern compilers and CPUs are too clever. A statement that a piece of code is slow can be incorrect, and a statement that it is fast can be incorrect as well. The only reliable route is to turn the claim into a measurement, which is why this chapter gives you the tools rather than a list of rules of thumb.

Measuring with std::chrono

C++ provides a portable, monotonic clock in the standard library: std::chrono::steady_clock. Unlike std::chrono::system_clock, the steady clock never jumps because of adjustments to the system time. Use steady_clock::now() before and after the code region and compute the difference. The following inline example measures the time taken to execute a trivial loop.

#include <chrono>
#include <iostream>

int main() {
    auto start = std::chrono::steady_clock::now();
    volatile int sum = 0; // prevent optimisation of the loop body
    for (int i = 0; i < 10'000'000; ++i) {
        sum += i;
    }
    auto end = std::chrono::steady_clock::now();
    auto elapsed = std::chrono::duration_cast<std::chrono::microseconds>(end - start);
    std::cout << "elapsed: " << elapsed.count() << " µs\n";
    return 0;
}

The program prints a single line such as elapsed: 12345 µs. The unit is a concrete, reproducible quantity that the test harness can match. Because steady_clock is monotonic, the measurement is not affected by clock adjustments, NTP updates, or daylight-saving changes. This makes it the preferred tool for micro-benchmarking code that runs for a short period.

For more reliable numbers, run the loop several times and record each measurement. Compute the median or the minimum value. The minimum discards noise from background activity, while the median reduces the impact of outliers. Warm up the code once before timing to let the processor reach its steady frequency and to populate caches. A typical benchmarking harness therefore performs a warm-up iteration, followed by a fixed number of timed iterations, and finally reports the best or median elapsed time.

The volatile in the timing loop is important. Without it the compiler can see that the loop has no observable effect and delete it entirely under the as-if rule, making the measured time zero. volatile forces the writes to happen, so the loop measures real work. The same trick is why micro-benchmarks accumulate into a volatile sink rather than returning a value that is never used.

A single measurement is not a number you can trust. Run the workload several times and report the spread, because a two-fold difference between runs is common under system noise.

The as-if rule and optimizer levels

The C++ as-if rule permits the compiler to transform any program as long as the observable behaviour is unchanged. Observable behaviour consists of the program’s side effects on volatile objects, file I/O, and the values returned from main. Therefore a build compiled with -O2 or -O3 can reorder statements, inline functions, or eliminate dead code, provided the resulting side effects match the source semantics. A build with -O0 performs almost no optimisation. It preserves the source order but does not represent the performance of a real-world binary. Benchmarking an unoptimised build therefore yields a number that the production binary will never exhibit. The meaningful comparison is always between two programs built with the same optimisation level.

Consider a function that adds two integers and returns the result. With -O0 the compiler emits a call to the function, a load of each argument, an addition, and a return. With -O2 the compiler can inline the function, keep the arguments in registers, and avoid the call entirely. The observable result, the returned sum, is identical, so the transformation is permitted. Benchmarks that report the speed of the -O0 version therefore mislead. They measure the cost of the extra call and the lack of register allocation, not the intrinsic cost of the algorithm.

Move vs copy

Moving a value transfers ownership of its resources without allocating or copying the underlying data. Copying, by contrast, must duplicate the resources. The instrumented Counter struct below records how many copy and move constructions occur. The book_example registration verifies the printed statistics. By examining the counters you can see that a move operation incurs far less work than a copy, especially when the underlying type manages heap memory or other expensive resources. This observation underlies the design of many standard library containers that prefer move over copy when they can.

#include <iostream>
#include <utility>

struct Counter {
    static int copies;
    static int moves;
    Counter() = default;
    Counter(const Counter&) { ++copies; }
    Counter(Counter&&) noexcept { ++moves; }
    Counter& operator=(const Counter&) = delete;
    Counter& operator=(Counter&&) = delete;
    ~Counter() = default;
};

int Counter::copies = 0;
int Counter::moves = 0;

int main() {
    Counter a;
    Counter b = a; // copy
    Counter c = std::move(a); // move
    std::cout << "copy: " << Counter::copies << " move: " << Counter::moves << "\n";
    return 0;
}

Running the program yields a line such as copy: 1 move: 1. The numbers confirm that the explicit copy and explicit move each invoke a single constructor, and that the default-constructed object does not contribute to the counts. If you replace the copy with another move, the copy counter stays at zero, showing the performance advantage of move semantics in realistic code. In larger containers, moving a std::vector merely swaps its internal pointer and size, while copying allocates new storage and copies each element, an order of magnitude more work.

Move operations are noexcept for the standard containers, and that single word unlocks a real optimisation. std::vector uses the move constructor during growth only when it is guaranteed not to throw. If the move can throw, the vector has to copy instead, to keep the strong exception guarantee. Marking your own types’ move constructors noexcept is therefore not ceremony. It is what lets vector move them during reallocation rather than copy.

See Chapter 5 for a detailed comparison of move versus copy costs.

Copy elision and RVO

When a function returns a prvalue of class type, the language permits the compiler to construct the result directly in the caller’s storage. This copy-elision eliminates both the copy and the move constructor calls. The classic case is the return value optimisation (RVO). The following example prints markers from the constructors and destructors. If elision occurs, only the constructor and destructor of the local object appear, and no copy or move messages are printed. This behaviour is guaranteed by the standard when the criteria for NRVO are met, and modern compilers perform it even at -O0.

#include <iostream>
#include <utility>

struct Marker {
    static int copies;
    static int moves;
    Marker() { std::cout << "ctor\n"; }
    Marker(const Marker&) { ++copies; std::cout << "copy\n"; }
    Marker(Marker&&) noexcept { ++moves; std::cout << "move\n"; }
    Marker& operator=(const Marker&) = delete;
    Marker& operator=(Marker&&) = delete;
    ~Marker() { std::cout << "dtor\n"; }
};

int Marker::copies = 0;
int Marker::moves = 0;

Marker make_marker() {
    Marker m; // ctor
    return m; // should be elided, no copy/move
}

int main() {
    Marker x = make_marker(); // elision expected
    (void)x;
    std::cerr << "copies=" << Marker::copies << " moves=" << Marker::moves << "\n";
    return 0;
}

The test harness expects the output to contain copies=0 moves=0. When the compiler performs RVO, the program’s output satisfies that expectation, demonstrating that the return did not incur any additional construction. If you deliberately disable copy-elision, for example by compiling with -fno-elide-constructors, the output changes to show a copy or move, which is useful for educational purposes but not representative of typical production builds.

Guaranteed copy elision, in effect since C++17, means a prvalue return does not even require the type to have a move constructor. A function that returns a prvalue of an immovable type still compiles and constructs the result in place. This is why returning a std::vector or a large struct by value is not just idiomatic but the fastest option. There is no copy and no move, only direct construction in the caller’s storage.

Chapter 18 showed that copy elision and RVO can remove all copy/move operations, making return‑by‑value the fastest way to deliver a result.

Container big-O review

Choosing the right container yields the highest performance gain in most programs. std::vector grows by amortised constant time. Each push_back is O(1) on average, but occasional reallocation costs O(n). Reserving capacity with reserve(n) eliminates those reallocations and therefore reduces the worst-case overhead. Associative containers differ. std::map provides ordered lookup in O(log n), while std::unordered_map offers average constant-time lookup, O(1), at the cost of higher memory usage and possible hash collisions. Understanding these complexities lets the programmer place the most expensive operations in the cheapest container. For example, building a large list of results is usually fastest with a vector that has been pre-reserved.

When a container holds objects that are expensive to move or copy, the cost of reallocation becomes significant. An optimisation is to store std::unique_ptr<T> in a vector and reserve enough space before filling it. This avoids repeated allocations of T and eliminates the need to move T objects during reallocation, because only the pointers are moved.

Big-O notation hides constant factors, which are real. An unordered_map is O(1) per lookup but has a large constant and high memory overhead, so for a handful of keys a linear scan of a small vector is faster. The rule is to choose by the shape of the workload, then confirm with a measurement, which brings the chapter’s central lesson back around. Choosing a container is a one-line change with large leverage, which is why it comes before micro-optimising a loop body.

Reading a hot loop

On Linux the profiler perf records CPU cycles, cache-miss events, and instruction retirements. On macOS the Instruments app provides similar metrics, including cache misses and allocations. When analysing a hot loop, look for a high proportion of cache-miss cycles, frequent allocations inside the loop body, and indirect calls such as virtual dispatch. Reducing cache misses involves improving data locality, for example by storing related objects contiguously in a vector or by using a struct-of-arrays layout. Eliminating allocations can be achieved with reserve or by reusing objects that are allocated once outside the loop. Virtual calls can be replaced by static polymorphism, by std::function_ref, or by inlining small call sites. The point is not to collect profiles but to find one or two dominant costs, because fixing the single hottest line usually beats optimising twenty small ones.

A typical workflow is:

  1. Run the program under perf record -g ./a.out or with Instruments’ time profiler.
  2. Identify the hottest functions from the flame graph.
  3. Drill into those functions to see which lines cause the most cache-miss or allocation events.
  4. Refactor the code to improve locality, pre-allocate storage, or replace virtual calls.
  5. Re-run the profiler to verify that the hot spots have diminished.

Try this

Predict whether calling reserve(n) on a std::vector<std::unique_ptr<Node>> that stores a binary tree will reduce the total runtime of a breadth-first construction loop. Measure the loop with std::chrono::steady_clock as shown earlier, print the elapsed time, and compare the two runs. Record the result as elapsed without reserve: X µs and elapsed with reserve: Y µs. The experiment demonstrates the real impact of pre-allocation on a realistic data-structure workload.

Appendix A: features and compilers

This appendix lists the C++ features the book uses. It shows the proposal that introduced or significantly changed each feature. It also shows the compiler support in the pinned toolchain (Clang 22.1.8, libc++). The required flag or header is listed. Finally it indicates whether the book compiles an example or marks it as a gap.

Compiled means the book ships a real example that builds and runs. Gap means the feature is taught prose-first and registered so it builds only with BOOK_ENABLE_GAPS=ON. A “Not yet deployable” callout names the supporting compiler if any.

FeatureProposalClang 22.1.8Flag / headerStatus
std::format / std::print / std::printlnP0645, P2093Yes<format>, <print>, -std=c++26Compiled
std::mdspanP0009Yes<mdspan>Compiled
std::numbers constantsP0631Yes<numbers>Compiled
ConceptsP0734Yes-std=c++20+Compiled
consteval, constinit, is_constant_evaluatedP1073, P1143, P0595Yes-std=c++20+Compiled
Transient constexpr allocationP0784Yes-std=c++20+Compiled
Class-type NTTPsP0732Yes-std=c++20+Compiled
User-defined literalsC++11Yes-std=c++20+Compiled
std::generatorP2502Yes<generator>Compiled
std::jthread / stop tokensP0660Yes<thread>, <stop_token>Compiled
std::atomic_refP0019Yes<atomic>Compiled
std::scoped_lockP0156Yes<mutex>Compiled
std::ranges / viewsP0896, P2325Yes<ranges>, <algorithm>Compiled
std::execution::parP0024No<execution>Prose only
Contracts ([[assert]], [[expects]], [[ensures]])P2900No (GCC 16)-fcontractsGap
std::execution (P2300 senders)P2300No<execution>Gap
std::linalgP1673No<linalg>Gap
Static reflectionP2996No (GCC 16 partial)<experimental/meta>Gap
<hazard_pointer> / <rcu>P2530, P2546No<hazard_pointer>, <rcu>Gap
Modules (import, export module)P1103Partial-std=c++26, BMIGap
Lifetime-safety analysisP1179Experimental-Xclang -fexperimental-lifetime-safetyCompiled (demos)

The Nix dev shell in flake.nix provides the pinned toolchain. The book compiles only verified examples. Newer features are taught with exact syntax and an honest callout.

Appendix B: coming from other languages

This appendix maps what you already know from C, Rust, Lisp, and Prolog onto the C++ you have now read. It is one page per language, and it uses the vocabulary of the chapter each idea belongs to.

From C

You knowIn C++ this is
char* strings with a NUL terminatorstd::string when owned, std::string_view when borrowed (ch10)
Arrays that decay to pointersstd::array, std::span, std::vector (ch11)
FILE* and manual fclosean RAII wrapper whose destructor releases the handle (ch05, ch28)
errno and sentinel returnsstd::expected and exceptions (ch09, ch28)
printf with unchecked format stringsstd::format and std::println, checked at compile time (ch10)
qsort with void* and a function pointerstd::ranges::sort with a comparator and projection (ch12)
Manual malloc/freestd::unique_ptr, std::vector, and the Rule of Zero (ch05, ch06)

C code works in C++ through the boundary discipline of chapter 28: extern "C" for linkage, spans and string views for C data, and RAII wrappers for C resources.

C++ retains C calling conventions, adds ownership, and replaces manual resource handling and unchecked formatting with RAII and compile‑time checked formatting.

From Rust

You knowIn C++ this is
Value semanticsvalue semantics, move semantics, and copy elision (ch05, ch29)
Ownershipa name owns a value. Ownership transfers on move (ch05)
The borrow checkerthe lifetime-safety profile, std::span, std::string_view, and -Wlifetime-safety (ch07)
Option<T>std::optional (ch03)
Result<T, E>std::expected (ch03, ch09)
Enums with datastd::variant plus std::visit (ch03)
Traitsconcepts and requires clauses (ch17)

Rust enforces ownership and lifetimes at compile time. C++ provides the same tools but relies on the lifetime‑safety analysis and disciplined APIs (see Chapter 7).

Both languages share the same mental model of ownership. The difference is where a violation is caught.

From Lisp

You knowIn C++ this is
Macros that expand sourcetemplates, which rewrite type patterns and are instantiated at compile time (ch16)
Compile-time evaluationconstexpr, consteval, and the compile-time execution model (ch18)
A domain-specific language evaluated at compile timethe constexpr SQL capstone (ch22)
Functions as datalambdas, std::function, and type erasure (ch14)
Recursive macros / term rewritingtemplate metaprogramming and pack expansion (ch16, ch21)

C++ templates and constexpr supply compile‑time code generation analogous to Lisp macros.

Lisp rewrites source text. C++ rewrites type patterns and evaluates a restricted subset of the language at compile time.

From Prolog

You knowIn C++ this is
Unificationtemplate argument deduction, which binds type parameters to concrete types (ch16)
Goal ordering / most specific ruleoverload resolution and concept subsumption (ch17)
A term with a tag and argumentsstd::variant plus std::visit (ch03)
Backtracking searchNot in the language. Express it explicitly with recursion or a search loop

The strongest analogy is deduction: Prolog unifies a query against rules, and C++ unifies a call against template patterns and selects the most specific viable match. Chapter 17 frames concept subsumption in exactly these terms.

The difference is control. Prolog searches for a solution and can backtrack. C++ resolves a call once at compile time and does not search at runtime. The shared idea is pattern matching against a set of rules, with the most specific rule winning.

Appendix C: Core Guidelines index

This appendix lists the C++ Core Guidelines rules referenced in the book, grouped by the chapter that teaches each rule. Rule IDs correspond to the CppCoreGuidelines version used during writing. When a rule appears in the text, its identifier is shown as (CG ) for easy lookup.

Chapter 2: values and functions

  • F.15: prefer simple and conventional ways of passing information.
  • F.20: for out-parameters, prefer return values over out-parameters.
  • F.21: to return multiple values, prefer returning a struct or tuple.
  • ES.25: declare an object const or constexpr unless you need to change its value.

Chapter 3: user-defined types

  • C.20: if you can avoid defining default operations, do.
  • C.2: use class if the class has an invariant.
  • C.2: use struct if the data members can vary independently.

Chapter 4: control flow

  • ES.28: use lambdas for complex initialization, especially of const variables.
  • ES.70: prefer a range-for-statement or a standard-library algorithm over a hand-written loop.

Chapter 5: ownership, move, and RAII

  • R.1: manage resources automatically using RAII.
  • R.3: a raw pointer or reference is never a resource owner.
  • R.10-R.11: avoid calling new and delete explicitly.

Chapter 6: smart pointers

  • R.20: use unique_ptr or shared_ptr to represent ownership.
  • R.22: use shared_ptr only to share ownership.

Chapter 7: lifetimes

  • Lifetime profile: never dereference a possibly-invalid pointer or iterator.
  • Lifetime profile: a pointer or reference to a local must not escape.

Chapter 8: argument passing

  • F.15: prefer simple and conventional ways of passing information.
  • F.20: for out-parameters, prefer return values over out-parameters.
  • F.21: to return multiple values, prefer returning a struct or tuple.
  • F.16-F.17: pass by value for small/cheap-to-move, const T& for big inputs.

Chapter 9: errors and contracts

  • E.1-E.16: error-handling rules, including throw when the function cannot do its job, and prefer std::expected for anticipated absence.
  • E.25: a function that cannot throw must be declared noexcept.

Chapter 11: containers

  • SL.con: standard-library container rules.
  • SL.con: prefer std::vector by default.

Chapter 12: algorithms

  • ES.70: prefer a range-for-statement or a standard-library algorithm over a hand-written loop.

Chapter 14: callables and type erasure

  • F.51-F.52: prefer a lambda or a function object for small callables.
  • F.51-F.52: prefer a regular function for a stateless callable.

Chapter 16-17: templates and concepts

  • T.1-T.65: template and generic-programming rules.
  • T.1-T.65: define concepts to express template constraints (T.10).

Chapter 21: reading legacy TMP

  • The rules for reading, not writing: write concepts, read SFINAE (T.1, T.10).

Chapters 25-26: concurrency

  • CP.1-CP.8: concurrency rules.
  • CP.1-CP.8: protect shared mutable state with a mutex or an atomic.
  • CP.1-CP.8: never access a non-atomic shared object without synchronization.

Chapter 29: fast is a specification

  • Perf.1-Perf.11: performance rules.
  • Perf.1-Perf.11: measure before optimizing (Perf.1).
  • Perf.1-Perf.11: avoid cheap micro-optimizations that do not show up in measurement.

This index is not exhaustive. The chapters cite the exact rule at the point of use. The intent of the book is that every rule it teaches is attributed, so a reader can chase the source of any guideline.

Appendix D: CMake

This appendix teaches CMake using the build setup of this book as the running example. CMake is a build‑system generator. It reads a description of the project and produces files for a native build tool. The book uses the Ninja generator. Ninja is a fast, low‑level build tool that CMake drives.

Configure, generate, build, test

CMake separates configuration from building. The configure step reads CMakeLists.txt files and records the choices you make. The generate step writes the Ninja files. The build step compiles the targets. The test step runs the registered tests.

The book drives all four steps with one command sequence:

cmake --preset dev
cmake --build build -j
ctest --test-dir build --output-on-failure

The first command configures and generates. The preset dev supplies the generator and the build directory. The second command builds every target with -j for parallel jobs. The third command runs every test and prints the output of failing tests. This sequence is the acceptance gate for the book. The script scripts/verify.sh runs it inside the Nix development shell.

The project file

The root CMakeLists.txt opens with a minimum version and a project name:

cmake_minimum_required(VERSION 3.30)
project(tour_cpp26 LANGUAGES CXX)

The version 3.30 guarantees that the features the book uses exist. The project declares that it uses only the C++ language. Three lines fix the language standard:

set(CMAKE_CXX_STANDARD 26)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)

The standard is C++26, strictly. CMAKE_CXX_STANDARD_REQUIRED ON refuses a compiler that cannot reach C++26. CMAKE_CXX_EXTENSIONS OFF forbids GNU extensions. The line set(CMAKE_EXPORT_COMPILE_COMMANDS ON) writes a compile_commands.json file that clangd and other tools consume.

Cache variables and options

CMake stores configuration values in a cache. The option command declares a boolean cache variable with a default:

option(BOOK_SANITIZE "Run example tests under AddressSanitizer and UndefinedBehaviorSanitizer" ON)
option(BOOK_ENABLE_GAPS "Build examples for C++26 features without shipping Clang support" OFF)
option(BOOK_WERROR "Turn compiler warnings into errors for book_example targets" ON)

BOOK_SANITIZE turns the sanitizers on for example tests. It defaults to ON. BOOK_ENABLE_GAPS builds examples for C++26 features that lack shipping Clang support. It defaults to OFF. BOOK_WERROR turns warnings into errors for book_example targets. It defaults to ON. You override a default on the command line with -D:

cmake --preset dev -DBOOK_ENABLE_GAPS=ON

The CMakePresets dev preset

A preset bundles configuration options under a name. The file CMakePresets.json defines the dev preset:

{
  "name": "dev",
  "generator": "Ninja",
  "binaryDir": "${sourceDir}/build",
  "cacheVariables": {
    "CMAKE_BUILD_TYPE": "RelWithDebInfo"
  }
}

The preset selects the Ninja generator. It places the build tree in the build directory beside the source. It sets the build type to RelWithDebInfo, which optimizes the code and keeps debug information. The preset also defines matching build and test presets, so cmake --build build -j and ctest --test-dir build work without extra flags.

Targets and subdirectories

A target is a named unit of work. The command add_executable creates an executable target from a source file. The command add_subdirectory descends into a child directory and reads its CMakeLists.txt. The file examples/CMakeLists.txt calls add_subdirectory for each chapter directory:

add_subdirectory(ch01)
add_subdirectory(ch02)

Each chapter directory holds one CMakeLists.txt that registers its examples. For example, examples/ch01/CMakeLists.txt contains:

book_example(ch01_hello.cpp EXPECT "hello, world")
book_demo(ch01_dangling.cpp)
book_demo(ch01_overflow.cpp)

The names book_example and book_demo are functions that the book defines. They are not CMake built‑ins.

The book_example function

The functions live in cmake/BookExample.cmake. The root CMakeLists.txt adds that directory to the module path and includes the file:

list(APPEND CMAKE_MODULE_PATH "${CMAKE_SOURCE_DIR}/cmake")
include(BookExample)

The book_example function registers a real, fully‑checked example. Its body is:

function(book_example file)
    cmake_parse_arguments(arg "" "EXPECT" "" ${ARGN})
    _book_add_target(t "${file}")
    target_compile_options(${t} PRIVATE ${BOOK_WARNING_FLAGS})
    if(BOOK_WERROR)
        target_compile_options(${t} PRIVATE -Werror)
    endif()
    if(BOOK_SANITIZE)
        target_compile_options(${t} PRIVATE -fsanitize=address -fsanitize=undefined)
        target_link_options(${t} PRIVATE -fsanitize=address -fsanitize=undefined)
    endif()
    add_test(NAME ${t}_run COMMAND ${t})
    if(arg_EXPECT)
        set_tests_properties(${t}_run PROPERTIES
            PASS_REGULAR_EXPRESSION "${arg_EXPECT}")
    endif()
endfunction()

The helper _book_add_target derives the target name from the file name and calls add_executable. The function attaches the warning flags, the -Werror flag, and the sanitizer flags. It registers a CTest test named <name>_run. When EXPECT is present, the test must match the given regular expression. The book_demo function registers a deliberately wrong example that compiles with warnings visible and no test. The book_gap function registers a C++26 feature target that builds only with BOOK_ENABLE_GAPS=ON.

The warning flags come from a list that includes the lifetime‑safety flag. The file probes the compiler with check_cxx_compiler_flag and check_cxx_source_compiles to discover which spelling of the lifetime‑safety analysis it accepts. It prefers -Wlifetime-safety and falls back to the experimental cc1 form.

Try this

Open cmake/BookExample.cmake, compare book_demo to book_example, then run cmake --preset dev -DBOOK_SANITIZE=OFF to verify the build succeeds without sanitizer flags.

Appendix E: CTest

This appendix teaches CTest, the test driver that runs the compiled examples of this book. CTest is part of CMake. It discovers, runs, and reports the tests that a project registers. The book uses CTest to exercise every book_example target and to verify its output.

Registering a test

A project enables testing with the enable_testing() command. The root CMakeLists.txt calls it before descending into the examples directory. A test is registered with add_test. The command names the test and gives the command to run:

add_test(NAME ch01_hello_run COMMAND ch01_hello)

The test ch01_hello_run runs the executable ch01_hello. CTest runs the executable in a separate process and reports whether it succeeded.

How book_example registers tests

The book_example function in cmake/BookExample.cmake registers one test for every example. The relevant lines are:

add_test(NAME ${t}_run COMMAND ${t})
if(arg_EXPECT)
    set_tests_properties(${t}_run PROPERTIES
        PASS_REGULAR_EXPRESSION "${arg_EXPECT}")
endif()

The target name t comes from the file name. The test name appends _run to it. So book_example(ch01_hello.cpp EXPECT "hello, world") produces the test ch01_hello_run. The _run suffix keeps the test name distinct from the executable name.

Expected output verification

A test passes when the program exits with status zero and its output satisfies the test properties. The property PASS_REGULAR_EXPRESSION holds a regular expression. When it is present, the test passes only if the program output contains a match for that expression.

A test fails when the program exits nonzero or when its output does not contain the expected regular expression. Both conditions are failures. The book relies on this rule to verify that an example prints what the text claims.

The EXPECT argument of book_example becomes the PASS_REGULAR_EXPRESSION. The book uses real regexes from the chapter files. A few examples:

book_example(ch01_hello.cpp EXPECT "hello, world")
book_example(ch02_quadratic.cpp EXPECT "roots: 2, 1")
book_example(ch03_point.cpp EXPECT "distance: 5.00")
book_example(ch03_parse_int.cpp EXPECT "value: 123|error at 2")

The last example shows a regex alternation. The vertical bar | means either alternative matches. The program prints value: 123 on success or error at 2 on failure, and the test accepts either.

Running the tests

The book runs the tests through the preset:

ctest --test-dir build --output-on-failure

The flag --test-dir build points CTest at the build directory that the dev preset created. The flag --output-on-failure prints the full output of a failing test. Without it, CTest shows only a summary line for each test. The dev test preset sets this behavior in CMakePresets.json, so the flag and the preset agree.

Naming, labels, and verbosity

Every test in this book is named after its example with a _run suffix. The naming makes a failing test easy to map back to its source file. The book does not use labels. A label groups tests under a name, and CTest can run only that group with ctest -L. The book has no need for grouping because every test is an independent example.

The default CTest output shows one line per test with a pass or fail result. The flag -V raises verbosity and prints each test command and its output. The flag -N lists the tests without running them. These flags help when you debug a single example.

Try this

Run ctest --test-dir build -N and list the tests that the book registers. Pick one example that uses an EXPECT regex, such as ch03_parse_int. Temporarily change its regex in the chapter CMakeLists.txt to a string the program never prints, rebuild, and run ctest --test-dir build --output-on-failure. Confirm that the test fails because the output does not contain the expected expression, then restore the original regex.

Appendix F: AddressSanitizer and UndefinedBehaviorSanitizer

This appendix teaches two runtime sanitizers that the book treats as members of the compiler committee. The address sanitizer (ASan) and the undefined‑behavior sanitizer (UBSan) detect errors while a program runs. They are part of the Nix development shell that flake.nix provides.

What ASan instruments

ASan instruments memory access. It detects a heap‑buffer‑overflow, which is a read or write past the end of a heap allocation. It detects a use‑after‑free, which is an access to memory that the program already released. It detects a memory leak, which is an allocation that the program never frees. ASan replaces the allocator and tracks every allocation and free. It checks each memory access against that bookkeeping.

The check happens at runtime, so the program must run to trigger a report. A program that never reaches the bad access stays silent. For this reason the book runs every example under the sanitizers, so a latent bug in an example surfaces during the test run.

What UBSan instruments

UBSan instruments operations whose behavior the standard leaves undefined. It reports signed integer overflow, which is an arithmetic result outside the representable range of a signed type. It reports shift overflow, which is a shift by a negative amount or by more than the width of the type. It reports a null‑pointer dereference, which is an access through a null pointer. UBSan inserts a runtime check before each offending operation and aborts as soon as the operation occurs.

UBSan and ASan complement each other. ASan catches memory errors. UBSan catches arithmetic and type errors. The book enables both together, so a single test run covers both classes of defect.

The flags

The sanitizers are enabled with compile and link flags. The compile flags instrument the code. The link flags attach the sanitizer runtime. The book uses both for each example:

-fsanitize=address -fsanitize=undefined

The same flags appear on the compile line and the link line. The address sanitizer and the undefined‑behavior sanitizer combine under one -fsanitize option. A thread sanitizer also exists, but on this toolchain it is Linux‑only. The book does not use it.

How BOOK_SANITIZE turns them on

The option BOOK_SANITIZE controls the sanitizers. It defaults to ON in the root CMakeLists.txt:

option(BOOK_SANITIZE "Run example tests under AddressSanitizer and UndefinedBehaviorSanitizer" ON)

The book_example function in cmake/BookExample.cmake applies the flags when the option is set:

if(BOOK_SANITIZE)
    target_compile_options(${t} PRIVATE -fsanitize=address -fsanitize=undefined)
    target_link_options(${t} PRIVATE -fsanitize=address -fsanitize=undefined)
endif()

Only book_example targets receive the sanitizer flags. The book_demo and book_gap targets do not. The sanitizers run under CTest, so a sanitizer report fails the test.

Runtime cost

The instrumentation stays in the finished binary. ASan roughly doubles runtime and memory use in typical programs. UBSan adds a check before each instrumented operation. For this reason the sanitizer builds are a test configuration, not the shipping binary. The book keeps the instrumented binaries in the test config only. The dev preset builds every example, and CTest runs each one under the sanitizers. The same source, built without BOOK_SANITIZE, produces the ordinary binary.

Reading an ASan report

The example ch01_overflow reads past the end of a vector to trigger a report. The report begins with a line that names the error class:

==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x6020000000fc
READ of size 4 at 0x6020000000fc thread T0
    #0 ... std::__1::__format::__create_format_arg ... format_arg_store.h:191
    ...
    #8 ... in main ch01_overflow.cpp:13

The first line names the error and the address. The second line describes the access, the size, and the thread. The stack trace follows. Each frame shows a call site. The final frame points into main at the source line of the offending read. The trace lets you walk from the top‑level call down to the exact statement that violated the rule.

Part of the Nix dev shell

The sanitizers are part of the Nix development shell. The file flake.nix provides the pinned toolchain, which includes LLVM 22.1.8. The shell sets CC=clang and CXX=clang++. The clang compiler ships the sanitizer runtimes, so the flags work without extra installation. Running nix develop and then the verify script builds every example and runs every test under the sanitizers.

The sanitizers are a test‑only configuration. They are not part of the shipping binary. The book enables them for the acceptance pipeline so that a memory error or an undefined operation in any example fails the build.

Try this

Write a small program that allocates an array with new[], reads one element past the end, and prints the value. Compile it with -fsanitize=address,undefined and run it. Confirm that ASan reports a heap‑buffer‑overflow with a stack trace that names your source line. Then write a program that adds two signed integers whose sum overflows, and confirm that UBSan reports the overflow while ASan stays silent.