A Tour of C++ for experienced programmers, as if C++26 is the only version that ever existed.
tourcpp book src ch24-modules.md
11 kB
Markdown
at main

Modules, the standard's direction #

What modules fix over #include #

The traditional include mechanism copies the text of a header file into each translation unit that names it. It also copies every macro definition that appears before the include directive. Because each translation unit receives its own copy of the header, the compiler cannot share work between units, which leads to long compile times for large projects.

A macro defined in one header can change the meaning of code that includes a later header. Thus, the order of header inclusion influences the program. Errors that originate inside a header are reported at the line that performed the include, making it hard to locate the source of the problem.

Modules replace this textual inclusion with a compiled interface. A module’s interface is built once. This build produces a binary module interface (BMI) file. The BMI contains only the declarations that the author chooses to export. Macros are not part of the BMI, so a macro defined in one translation unit cannot affect a module that imports it. Errors that arise inside a module are reported inside the module file itself, giving a clear location. The build system can reuse the BMI for every importer, which reduces compile time dramatically for projects that import the same module many times. In short, modules give isolation, order independence, and faster incremental builds.

import std #

The C++ standard library can be treated as a single module. The statement

import std;

brings every name from the library into the program. No header file needs to be included. The compiler reads the BMI for the standard library module instead of opening dozens of header files. Unfortunately the pinned toolchain (Clang 22.1.8) does not recognise the import keyword, so a file that contains the line above fails to compile. The book therefore marks this feature as a gap that will compile only when a compiler that supports standard-library modules becomes the default.

Not yet deployable. The current Clang 22.1.8 toolchain does not understand import. The example file examples/ch24/import_std.cpp is registered as a gap. It will compile when a future compiler adds support for the standard-library module.

Named modules #

A named module defines a logical unit of code that other translation units can import. The first file that declares the module is the interface unit. It contains the export module declaration followed by the declarations that the author wishes to make visible.

// examples/ch24/mymod.cppm
export module mymod;
export int add(int a, int b);

Implementation units provide the definitions for the exported declarations. An implementation unit begins with module <name> and then defines the functions.

// examples/ch24/mymod_impl.cpp
module mymod; // implementation unit
int add(int a, int b) { return a + b; }

When a program wishes to use the module it writes an import statement.

// examples/ch24/main_using_mymod.cpp
import mymod;
int main() { return add(2, 3); }

The import causes the compiler to load the BMI that was created from mymod.cppm. The linker then resolves the definition that lives in mymod_impl.cpp. The two source files together form a complete module.

Modules can be split into partitions when a large module needs many source files. A partition is declared with a colon after the module name.

export module mymod:part;
export int sub(int);

The corresponding implementation unit uses the same module mymod:part; header. Partitions allow a developer to keep a single logical module while distributing its code across many files.

The example files in examples/ch24/ follow this pattern. They are registered as a gap because the current compiler does not accept the module keyword.

Naming modules follows a simple convention: use lower‑case identifiers that reflect the library’s purpose, avoid mixed‑case or digits, and keep the name stable across versions. Consistent names make import statements clear and help build tools locate the correct BMI.

{{#include ../../examples/ch24/mymod.cppm}}
{{#include ../../examples/ch24/mymod_impl.cpp}}
{{#include ../../examples/ch24/main_using_mymod.cpp}}

Build integration and testing #

Testing modules follows the same pattern as testing header-only code. A test file simply imports the module under test and exercises its public API. Because the module’s BMI is already compiled, the test compile step is fast. Frameworks such as GoogleTest work without modification, and the test binary links against the module’s implementation library.

Adoption reality #

Adopting modules in a legacy code base must be incremental. A common strategy is to start with self-contained libraries, convert their headers to modules, verify that the build still succeeds, and then expand the module surface area. Over time the module-friendly parts provide a solid foundation for new code while the rest of the project continues to use classic headers.

The chapter records the module examples as gaps using the book_gap macro because the current toolchain does not support the module syntax. When a compiler that supports modules is used, the same CMake configuration will compile the interface, generate the BMI, and link the implementation automatically.

Module build integration details #

A CMake target that represents a module interface is created with the source file that contains the export module declaration. CMake adds the -fmodule-file= flag automatically for any target that lists the module as a dependency. The generated BMI file has the extension .pcm on Clang and is placed in the build directory alongside other compiled objects. Importing targets read the BMI directly, which avoids reparsing the source file.

When the interface source changes, CMake rebuilds only the BMI and any dependents that import the module. The implementation unit is compiled as a static library that provides the definitions for the exported symbols. The static library is linked into any executable or library that imports the module. Because the interface and implementation are separate, developers can modify the implementation without triggering a rebuild of the interface, further reducing incremental build time.

CMake also propagates include directories from the module target to its dependents, so that headers used inside the module are found without additional target_include_directories calls. This behavior mirrors the way the compiler handles the standard library module.

The book marks these examples as gaps, but the same CMake patterns work unchanged on a compiler that implements modules. The author encourages readers to try the conversion on a supporting toolchain and compare build metrics.

Mixing modules and headers #

A module can import a classic header file. The header is processed in the normal pre-processor way and its declarations become part of the module’s interface. The module can then re-export those names if desired.

export module mymod;
import <vector>;          // import a header
export using std::vector; // re-export the type

The opposite direction is not permitted. A header file cannot contain an import statement because the pre-processor runs before the module system and does not recognise the keyword. Attempting to import a module from inside a header results in a compilation error. Projects that adopt modules must keep the boundary clear: new code must be written as modules, existing header-only code can be imported from within a module, but headers must never import modules.

#embed #

The pre-processor provides a directive #embed that inserts the raw bytes of a file as a constant array at compile time. The syntax is straightforward.

const unsigned char logo[] = {
    #embed "assets/logo.png"
};

The majority of production code still relies on the header-include model. Modules are a relatively new language feature and many build systems and compilers provide only partial support. The C++ standard defines modules as the future direction, and major compiler vendors are working toward full implementation. Readers will encounter both models in the wild. Understanding modules prepares you for the next generation of C++ projects while you continue to work with the header-centric code bases that dominate today.

Developers can measure the build impact by compiling a representative set of files with and without modules. Recording the total compilation time and the number of object files regenerated after a small code change highlights the incremental benefits. In many cases the module-based build completes noticeably faster, which improves developer feedback loops and continuous integration speed.

Module benefits #

Modules isolate code by compiling a clean interface that contains only exported declarations. This isolation removes the impact of macros defined in other translation units.

Because the interface is compiled once, the compiler can reuse the binary module interface (BMI) for every importer. This reuse reduces parsing work and speeds up incremental builds.

The module system also provides order‑independence. Import statements do not depend on the order of header inclusion, so changes in one header do not cause unrelated recompilation.

Partitions allow a large module to be split across several source files while keeping a single logical name. Each partition is declared with a colon after the module name and can be compiled separately. Partitions enable developers to group related functionality while preserving a single import statement for the whole module, simplifying dependency management.

The book includes a page that explains the macro‑isolation argument and the use of partitions in detail.

Try this #

Convert the Vec2 header/source pair from chapter 23 into a module pair (an interface unit and an implementation unit). Keep the main program’s behaviour unchanged. Record any changes needed in the CMake configuration and note that this conversion can be compiled and compared on a compiler that supports modules.

  • Write vec2.cppm as the interface unit exporting the Vec2 type and its functions.
  • Write vec2_impl.cpp as the implementation unit defining the functions.
  • Update the example’s CMake target to use book_gap for these files, because the current toolchain does not support module syntax.
  • Verify that the program compiles and runs on a supporting compiler and observe the build‑time impact.
  • Document the required CMake changes and any differences in build output.

This exercise demonstrates module isolation and the build‑system integration steps.