Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

User-defined types: structs, sums, and expectations

Structs are product types

In C++ a struct groups a fixed set of named fields. The compiler automatically provides a default constructor, a copy constructor, a move constructor, and a trivial destructor when the fields themselves support those operations. This mirrors the mathematical notion of a product: a value of the struct type contains one value of each field.

struct point {
    double x;
    double y;
};

The declaration above yields a type that can be constructed with brace initialisation: point{1.0, 2.0}. No user-written constructor is required, which satisfies the Core Guidelines recommendation to prefer aggregates (C.20). The fields are public by default. The type behaves like a plain data carrier, unlike the historic C struct, which cannot contain member functions or const qualifiers.

Member functions can be added without sacrificing the aggregate property. A read-only member function is marked const so the compiler forbids it from modifying the object (C.2). The point example below adds double distance(const point&) const, which computes the Euclidean distance without changing *this and can therefore be called on a const point.

An aggregate also supports designated initialisers. The syntax point{.x = 1.0, .y = 2.0} names each field explicitly. The compiler checks the order and the names, so a typo in a field name fails to compile. Designated initialisers make the intent of each value clear at the call site. They also survive a change in field order without silently swapping values. This form is the clearest way to build a small data carrier.

The point type with its distance function is a complete example. The test verifies that the distance between two points is correct.

#include <print>
#include <cmath>

struct point {
    double x;
    double y;
    double distance(const point& other) const {
        double dx = x - other.x;
        double dy = y - other.y;
        return std::hypot(dx, dy);
    }
};

int main() {
    point p1{0.0, 0.0};
    point p2{3.0, 4.0};
    double d = p1.distance(p2);
    std::println("distance: {:.2f}", d);
    return 0;
}

Invariants and the small private we allow

A struct needs to enforce a relationship between its fields: an invariant. The book avoids class-based object orientation, yet a narrow use of private together with a public accessor is acceptable when the invariant cannot be expressed by the type system alone. For example, a normalized 3-D vector must always have length 1:

struct vec3 {
private:
    double x, y, z;
    vec3(double a, double b, double c) : x(a), y(b), z(c) {}
public:
    static vec3 make(double a, double b, double c) {
        double len = std::hypot(a, b, c);
        return vec3{a / len, b / len, c / len};
    }
    double length() const { return std::hypot(x, y, z); }
};

The private section prevents accidental mutation that can break the invariant, while the static factory make guarantees a correctly normalised instance. This follows Core Guidelines C.21 to keep data encapsulation minimal and prefer plain functions over heavy OO machinery. The invariant, that a vec3 constructed via make has length 1, is enforced by the factory. Direct use of the private constructor violates the contract, and the factory also checks for a zero‑length input and rejects it, avoiding a NaN result.

Because the type system cannot express this invariant, the responsibility rests with the factory and disciplined callers. This trades a small runtime check for the complexity of a large class hierarchy.

Defaulted comparison and the spaceship

C++20 introduced the three-way comparison operator <=>, commonly called the spaceship. When a struct’s fields already support equality and ordering, the compiler can generate all six relational operators automatically:

struct point {
    double x;
    double y;
    auto operator<=>(const point&) const = default;
};

Before C++20 each relational operator had to be written manually, a source of boilerplate and possible inconsistencies. The defaulted spaceship satisfies the Core Guidelines rule that the compiler must generate “obvious” functions (C.10, C.87). The generated operators perform a lexicographic comparison of the fields in declaration order, mirroring the behaviour of std::tie.

The spaceship is the last piece the point type needs. With it, the type supports equality, ordering, and structured bindings, all without a hand-written operator. The compiler generates the comparisons from the fields, and the fields are the data. This is the Rule of Zero applied to comparisons: the type declares its intent (= default), and the compiler does the work.

Structured bindings let a caller unpack a point into its fields in one statement. The form auto [px, py] = p binds px to p.x and py to p.y. The binding respects the declaration order of the fields, which matches the order that the spaceship compares. So the ordering of the fields has two consequences: it fixes the order of comparison and the order of the bindings. A reader who knows the field order can predict both. This consistency is why the book keeps field order stable across the type.

The spaceship also reports the strength of the comparison. The defaulted operator returns a comparison category, and the compiler derives it from the field types. For a point of two double fields the category is std::partial_ordering, because floating-point values do not always compare as equal to themselves. The caller rarely names this category directly, yet it governs how the generated operators behave. This detail matters when a struct mixes fields of different comparison strength.

std::variant as a sum type

A sum type represents a value that is exactly one of several alternatives. In C++ the standard library provides std::variant for this purpose. It stores a discriminated union together with a runtime tag, guaranteeing safe access. The expression-tree example below demonstrates a binary addition tree built from numbers and nested operations:

#include <variant>
#include <memory>
#include <iostream>

struct Node;
struct Number { double value; };
struct BinaryOp {
    char op;
    std::unique_ptr<Node> left;
    std::unique_ptr<Node> right;
};
struct Node {
    std::variant<Number, BinaryOp> expr;
    double eval() const {
        struct Visitor {
            double operator()(const Number& n) const { return n.value; }
            double operator()(const BinaryOp& b) const {
                double l = b.left->eval();
                double r = b.right->eval();
                return b.op == '+' ? l + r : 0.0;
            }
        };
        return std::visit(Visitor{}, expr);
    }
};

int main() {
    // Build (1 + 2) + 3 = 6
    auto leaf1 = std::make_unique<Node>(Node{Number{1}});
    auto leaf2 = std::make_unique<Node>(Node{Number{2}});
    auto inner = std::make_unique<Node>(Node{BinaryOp{'+', std::move(leaf1), std::move(leaf2)}});
    auto root  = std::make_unique<Node>(Node{BinaryOp{'+', std::move(inner), std::make_unique<Node>(Node{Number{3}})}});
    std::println("result: {}", root->eval());
}

std::visit dispatches to the appropriate overload based on the active alternative. The helper Visitor aggregates the lambdas (a pattern often called overloaded) and makes the code concise. Compared with a traditional C union, std::variant carries a tag, eliminating the undefined behaviour that arises from reading the wrong member. Rust’s enum offers a similar safety guarantee. The C++ version fits naturally into the existing type system without requiring a separate language construct.

The expression tree is a Prolog term in C++ clothing. A Node is either a Number (a leaf) or a BinaryOp (a compound term with two subterms). The eval function is the interpreter: it walks the term and reduces it to a value. The visit call is the case dispatch, and the Visitor struct is the set of clauses, one per alternative. This is the pattern the book returns to in chapter 22, where a compile-time SQL parser builds a similar term and evaluates it at translation time.

The full expression-tree example is a complete program. The test verifies that the tree evaluates to the correct value.

#include <variant>
#include <memory>
#include <iostream>

struct Node;
struct Number { double value; };
struct BinaryOp {
    char op;
    std::unique_ptr<Node> left;
    std::unique_ptr<Node> right;
};
struct Node {
    std::variant<Number, BinaryOp> expr;
    double eval() const {
        struct Visitor {
            double operator()(const Number& n) const { return n.value; }
            double operator()(const BinaryOp& b) const {
                double l = b.left->eval();
                double r = b.right->eval();
                return b.op == '+' ? l + r : 0.0;
            }
        };
        return std::visit(Visitor{}, expr);
    }
};

int main() {
    // Build (1 + 2) + 3 = 6
    auto leaf1 = std::make_unique<Node>(Node{Number{1}});
    auto leaf2 = std::make_unique<Node>(Node{Number{2}});
    auto inner = std::make_unique<Node>(Node{BinaryOp{'+', std::move(leaf1), std::move(leaf2)}});
    auto root  = std::make_unique<Node>(Node{BinaryOp{'+', std::move(inner), std::make_unique<Node>(Node{Number{3}})}});
    std::println("result: {}", root->eval());
}

std::optional for the maybe

A function that can return a value or nothing uses std::optional<T>. An optional either holds a T or is empty. It replaces the raw-pointer-returns-null idiom, which a caller can forget to check. The value() accessor throws std::bad_optional_access if the object is empty, and operator* is undefined for an empty optional, so the type makes the empty state explicit in the API. This pattern aligns with the Core Guidelines principle that error-prone pointer use must be avoided (F.4).

An optional carries no information about why the value is absent. It answers the question “is there a value?” and nothing more. When the caller needs to know why the operation failed, std::expected is the right type, because it carries an error description alongside the absence.

The choice between optional and expected is a question of what the caller must know. Use optional when the absence is a normal state, not an error. A map lookup that finds no key is a good fit, because the caller can proceed without the value. Use expected when the absence is a failure that the caller must handle or report. A file read that fails needs a reason, so the caller can log it or retry. The two types share the same shape, yet they carry different meaning. The book states the rule plainly: absence without a reason is optional, absence with a reason is expected.

std::expected for fallible computations

When a computation can fail with a recoverable error, std::expected<T, E> conveys either a successful result of type T or an error description of type E. The parse-integer example returns an int on success or a parse_error struct describing the position of the first non-digit character.

#include <expected>
struct parse_error { int position; std::string_view message; };
std::expected<int, parse_error> parse_int(std::string_view sv);

The caller inspects the returned object with if (result) and branches on success or failure. This approach is preferable to exceptions for anticipated failure such as malformed user input, because the control flow is explicit and does not incur the cost of stack unwinding. It also improves on std::optional by providing diagnostic information: the error type can carry a message, an error code, or any richer context the application needs. Chapter 9 expands on the trade-offs between exceptions, expected, and optional.

The parse-integer example is a complete program. The test verifies that a valid input produces the correct integer and an invalid input produces the error message.

The expected type also supports composition. A caller can chain operations so that a failure stops the chain early. The member function and_then applies a follow-up function only when the object holds a value. The member function or_else supplies a fallback when the object holds an error. These operations keep the control flow in the value domain instead of in a set of nested checks. The result reads like a pipeline of steps, and the error path stays visible in the type. This is the same spirit as the visit dispatch in the expression tree: the shape of the type drives the flow of the code.

The error type E is a full type, not a fixed string. A program can define an error that carries a code and a message, or a nested error from a lower layer. The caller can inspect the error and decide how to recover. This flexibility is what separates expected from a bare bool return. A bool tells the caller that something failed, while an expected tells the caller what failed and where. The extra information is the reason the book reaches for expected over a status flag.

#include <print>
#include <string_view>
#include <expected>

struct parse_error {
    int position;
    std::string_view message;
};

std::expected<int, parse_error> parse_int(std::string_view sv) {
    int result = 0;
    int pos = 0;
    for (char c : sv) {
        if (c >= '0' && c <= '9') {
            result = result * 10 + (c - '0');
            ++pos;
        } else {
            return std::unexpected(parse_error{pos, "non-digit character"});
        }
    }
    return result;
}

int main() {
    auto good = parse_int("123");
    if (good) {
        std::println("value: {}", *good);
    } else {
        std::println("error at {}", good.error().position);
    }
    auto bad = parse_int("12a3");
    if (bad) {
        std::println("value: {}", *bad);
    } else {
        std::println("error at {}", bad.error().position);
    }
    return 0;
}

Try this

Extend the point example from the product-type section by adding a defaulted three-way comparison (operator<=>). Create three point objects, store them in a std::vector<point>, and call std::sort. State in one sentence which comparison the sort algorithm used.