Headers, #include, and organizing programs
The build model everyone uses
Modern C++ programs are built from many translation units. A translation unit is the result of a single source file after the preprocessor has run. The compiler translates each unit independently to an object file and a linker later combines all object files into an executable or a library.
The preprocessor step is a literal text substitution. When the source contains a line such as
#include "vec2.hpp"
the preprocessor opens vec2.hpp, copies every line into the source, and then continues scanning the combined text. No compilation happens before the whole file has been assembled. All macro expansion, conditional compilation, and header inclusion happen at this stage. The resulting file is what the compiler sees as one translation unit.
Because the process is entirely textual, a header can be included many times, from many different source files, and even through different relative paths. If a header is pulled in twice the same name appears twice in the same translation unit, which is illegal unless the header is protected.
Modules, introduced in the next chapter, provide a different mechanism that bypasses textual pasting. For now the dominant reality is the include-paste model described above.
Separate compilation has a practical payoff. When one source file changes, only that unit recompiles, and the linker recombines it with the unchanged object files. This is why a large project rebuilds quickly after a single edit instead of recompiling everything. The compiler sees each unit in isolation, so it learns names from headers.
Declarations versus definitions
A declaration introduces a name to the compiler. A definition supplies the complete entity that the name represents.
Typical examples:
// Declaration – tells the compiler that a function exists.
double length(const Vec2& v);
// Definition – provides the body that implements the function.
double length(const Vec2& v) { return std::sqrt(v.x*v.x + v.y*v.y); }
Headers normally contain declarations. Source files (.cpp) contain the matching definitions.
To call a function, the compiler needs only its declaration: the name, the parameter types, and the return type. The body can live in a different translation unit, compiled separately and found by the linker. For a class, however, the full definition must be visible wherever the class is used by value or its members are accessed, which is why class definitions live in headers.
Templates and inline functions are an exception: the definition must be visible to every translation unit that uses them, so the definition itself lives in a header. The same rule applies to constexpr variables: their definition is required in each unit that odr‑uses the variable.
Include guards and #pragma once
If the same header is included twice, the preprocessor copies the same declarations and definitions twice, producing duplicate symbols and compile‑time errors. Two mechanisms prevent this duplication.
Classical include guard
#ifndef VEC2_HPP // If VEC2_HPP not defined …
#define VEC2_HPP // … define it and process the file.
... header contents …
#endif // VEC2_HPP // End of guarded region.
The first inclusion defines the macro. Subsequent inclusions see the macro already defined and skip the whole file.
#pragma once
Many compilers support the single directive
#pragma once
placed at the top of a header it guarantees the file is processed at most once per translation unit. The effect is identical to the classic guard but requires fewer lines and avoids accidental macro name clashes.
Both forms are accepted by the book’s build. The examples show the guard form for portability.
The One Definition Rule
The One Definition Rule (ODR) states that every non‑inline function, variable, class, or template specialization must have exactly one definition in the entire program.
If two translation units each contain a definition of the same non‑inline function, the linker reports a multiple‑definition error.
Headers are the root cause of many ODR violations. A header that contains a full definition of a normal function will be copied into every translation unit that includes it, creating multiple definitions. That is why most functions belong in source files, while only declarations appear in headers.
The ODR also applies to types: two different definitions of a class with the same name break the rule, even if the definitions are textually identical. The linker cannot merge class definitions. The program must contain a single authoritative definition.
A type-level ODR violation is the subtlest: differing class layouts across units cause memory corruption, so define each class once in a header and include it everywhere. Likewise, placing a non‑inline function definition in a header creates multiple definitions. Move the definition to a source file and keep only the declaration in the header.
inline as the ODR valve
The inline specifier deliberately relaxes the ODR for functions and variables.
- an inline function can be defined in any number of translation units, provided each definition is identical after preprocessing.
- an inline variable (C++17 onward) follows the same rule.
When a header defines an inline function or variable, each translation unit that includes the header gets its own copy. The linker discards the duplicates and keeps a single entity.
Because the definitions are required to be identical, the compiler can safely replace a call with the function body (inline expansion) or keep a single out‑of‑line copy if necessary.
The book uses inline constexpr double pi = 3.141592653589793;. The definition must be identical in every translation unit, otherwise behavior is undefined.
Linkage and anonymous namespaces
Linkage determines whether a name is visible across translation units.
| Linkage type | Visibility |
|---|---|
| external | visible to the linker. The name can be used from any translation unit. |
| internal | visible only inside the translation unit where it is defined. |
The static keyword on a namespace‑scope variable gives it internal linkage (deprecated for functions, but still valid). A const variable at namespace scope also has internal linkage unless explicitly marked extern.
An anonymous namespace provides internal linkage for every name declared inside it:
namespace {
int helper() { return 42; } // internal linkage
}
The compiler assigns a unique mangled name to each translation unit, guaranteeing that the helper does not clash with a helper in another file. This technique is useful for implementation‑detail functions that must not appear in the public symbol table.
The choice between internal and external linkage is an interface decision. Names with external linkage are part of the program’s symbol table and can collide with other units. Names with internal linkage are private to their unit, so two units can each define a detail helper with the same name without conflict. This is what makes anonymous namespaces the standard way to hide implementation helpers.
The multi‑file example
The following three files implement a tiny 2‑D vector library. The header declares the type and its interface, the source file defines the functions, and main.cpp uses the library.
#ifndef VEC2_HPP
#define VEC2_HPP
#include <cmath>
// Inline constant – visible to every translation unit.
inline constexpr double pi = 3.14159265358979323846;
struct Vec2 {
double x{};
double y{};
// Defaulted constructors.
Vec2() = default;
Vec2(double x_, double y_);
// Returns the squared length – useful for comparisons.
double length_sq() const;
// Returns the Euclidean length.
double length() const;
// Adds another vector to this one.
void add(const Vec2& other);
};
#endif // VEC2_HPP
#include "vec2.hpp"
Vec2::Vec2(double x_, double y_) : x(x_), y(y_) {}
double Vec2::length_sq() const {
return x * x + y * y;
}
double Vec2::length() const {
return std::sqrt(length_sq());
}
void Vec2::add(const Vec2& other) {
x += other.x;
y += other.y;
}
#include "vec2.hpp"
#include <iostream>
int main() {
Vec2 v(3.0, 4.0);
std::cout << "length " << v.length() << '\n';
return 0;
}
The directory also contains a CMakeLists.txt that builds the program as a raw executable and registers a test that checks the printed output:
add_executable(ch23_two_files main.cpp vec2.cpp)
add_test(NAME ch23_two_files_run COMMAND ch23_two_files)
set_tests_properties(ch23_two_files_run PROPERTIES
PASS_REGULAR_EXPRESSION "length 5"
)
The build command that a reader runs from the repository root is
cmake --preset dev -S . -B build && cmake --build build --target ch23_two_files
Running the test with ctest --test-dir build confirms that the program prints the expected length of a 3‑4‑5 right triangle.
Compiling the three files separately is instructive. vec2.cpp compiles vec2.hpp, so any mismatch between declarations and definitions fails here. main.cpp compiles vec2.hpp again, and the linker merges the two object files. The header is therefore compiled twice, once per unit, which is exactly why it must be self-contained and guarded, and why editing it forces every including unit to rebuild.
Header hygiene
A self‑contained header compiles on its own. That means it includes every header it needs, and nothing else.
Never rely on a transitive include from another header. If vec2.hpp needs <cmath> it must include it directly, even if another header already includes <cmath>. This prevents surprising compile errors when the header is used elsewhere.
Include what you use is the guiding principle.
If a header only uses a forward declaration of a class, it must forward‑declare rather than include the full definition. This reduces compile‑time dependencies and avoids unnecessary recompilation when unrelated headers change.
The example library follows these rules:
vec2.hppincludes only<cmath>because the header needs thestd::sqrtdeclaration for the inlinelength()definition.- the source file
vec2.cppincludes the same header to ensure the declarations match the definitions. main.cppincludes onlyvec2.hppand the standard<iostream>for output.
Every header must be compiled in isolation to detect missing includes, and forward declarations replace full includes when only pointers or references are used.
Try this
Split the expression‑tree variant example from chapter 3 into two files: a header that declares the Expr type and its visitor, and a source file that defines the evaluation function. Keep the program’s behaviour unchanged and make sure the build still passes the existing test.