[0:20 — 0:20]
Hi everyone, my name is Laurent Carlier. I'm a freelance consultant from Belgium, and I write
embedded and low-level C++ for a living. Today, we're going to talk about templates.
Goals of this presentation
Take home something you can use
Template enables type safety
Template is performant
Template enable generic programming
[0:50 — 1:10]
Before we start, let me tell you what I want you to get out of this talk.
My first goal is that you take home something you can use: by the end of this presentation, you
should be able to go and write your own templates.
And to convince you that it's worth doing, I'm going to show you three concrete reasons along the
way: templates give you type safety, they are performant, and they open the door to generic
programming.
Goals of this presentation
Template doesn't need to be complicated
[0:30 — 1:40]
The second goal is to show you that templates don't have to be complicated. With just a handful
of features, you can already write efficient and performant code.
And by the way, I gave this presentation to my cat, and she understood everything. So I'm sure
you will too.
Presentation philosophy
Teach by examples
Explain the feature used in the examples
Warn about common pitfalls
[1:25 — 3:05]
Templates are a huge topic, and you could easily spend days on them. So, given the goals of this
talk and the time I have, I settled on the following philosophy.
I like to keep things simple, so I'm going to show you a lot of examples. And at the end of the
presentation, we'll look at three examples taken from production code.
To make sure those examples make sense to you, I'll explain, from the basics, every feature they
use. Because of the time constraint, I won't be able to go further than that.
And of course, this is a Back to Basics talk, so I'll also point out the common mistakes people
make with these features.
After this talk, I'd really encourage you to watch Nicolai Josuttis's recording. He gave a talk
with the same title, and he goes into much more detail than I will today.
⏰ Agenda of today ⏰
Syntax & function templates
Implicit requirements
Template argument deduction
Variadic templates & perfect forwarding
Class templates
3 production code examples
[0:25 — 3:30]
This slide gives an overview of what we will cover today. We will start with syntax and function
templates, then move on to implicit requirements, template argument deduction, variadic
templates and perfect forwarding, class templates, and finally, three production code examples.
Template syntax
A template defines a family of functions , classes , type aliases ,
variables or concepts .
template< [parameter-list] >
recipe
parameter-list can be:
type parameter e.g. std::string
non type parameter (aka NTTP) e.g. enum value, constexpr value
parameter pack using ellipsis ...
template parameter e.g. template<typename> class
recipe can be:
function definition
a class/struct definition
a type alias definition
variable definition
concept definition
[0:45 — 4:15]
A template comes in exactly 5 forms: class templates, function templates, alias templates,
variable templates, and concepts.
The syntax works as follows. You write the keyword `template`, followed by angle brackets
containing the parameter list, and then the declaration of the entity you are defining.
In between the angle brackets, you list the template parameters, which can be type parameters,
non-type parameters, parameter packs, or template parameters.
The recipe then contains the declaration of the form that you are defining.
Function template
Defines a family of functions
template<typename Swappable>
void mySwap(Swappable& a, Swappable& b) {
Swappable tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
💡 Give good name to template parameters
⚠️ The compiler does not produce any code if the template is not instantiated
[1:10 — 5:25]
Let's define our first family. Here I'm showing a function template that defines a family of swap
functions for different types.
To tell the compiler that we expect a type parameter, we use the keyword typename followed by a
template parameter identifier. Notice that it is not T.
Then we can use that identifier to declare function parameters in our function signature or we
can also use that identifier to declare an instance of that type in our function.
This really looks like a function but the difference with a normal function is that some types
are not known at the time of writing the function. Those are the template parameters.
Because the types are not known, the compiler does not produce any code at this stage. We need to
instantiate the template.
Function template
Defines a family of functions
template<typename Swappable>
void mySwap(Swappable& a, Swappable& b) {
Swappable tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
int a = 5, b = 10;
mySwap(a, b);
assert(a == 10 && b == 5);
float c = 3.14f, d = 2.71f;
mySwap(c, d);
assert(c == 2.71f && d == 3.14f);
⚠️ Calling the function with 2 different types fails
float f = 3.14f;
int i = 42;
mySwap(i, f); // ❌ Does not compile
[1:05 — 6:30]
In order to instantiate the template, we need to call it like we would call a regular function.
For instance, if I declare 2 ints a and b, I can call mySwap with a and b.
Similarly, if I declare 2 floats c and d, I can call mySwap with c and d.
Notice that in both cases I didn't have to specify the template parameter. This is because we
have a feature called template argument deduction and we are going to come back on it later.
Now if I call mySwap with 2 variables of 2 different types, the code will fail to compile. This
shows that the template enforces type consistency for its parameters hence enabling type safety.
Function template
Defines a family of functions 🔎
cppinsights.io/s/b05df164
Source:
template<typename Swappable>
void mySwap(Swappable& a, Swappable& b) {
Swappable tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
int a = 5, b = 10;
mySwap(a, b);
float c = 3.14f, d = 2.71f;
mySwap(c, d);
Insight:
template<>
void mySwap<int>(int& a, int& b)
{
int tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
template<>
void mySwap<float>(float& a, float& b)
{
float tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
⚠️ The compiler generates code for each instantiation.
https://cppinsights.io/s/b05df164
[0:50 — 7:20]
Let's see exactly the impact of template instantiation on our program. To do that I'm going to
use C++ Insights. It is a fantastic tool developed by Andreas that runs a Clang front-end and
reveals the code the compiler produces for templates.
We can see that for each instantiation of mySwap, the compiler generates a separate, fully-typed
function. This can have an impact on the binary size. The more instantiations with different
types, the bigger the code size.
However the optimizer is there to help you and the outcome is the same as if you would have
written the function separately.
Function template
How are function templates compiled?
Two-phase name lookup
template<typename Swappable>
void mySwap(Swappable& a, Swappable& b) {
Swappable tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
int a = 5, b = 10;
mySwap(a, b);
First, the compiler checks for syntax and resolve non-dependent names identifier.
Second the template argument is substituted in for the template parameter
⚠️ The code can compile until a function template is called
[1:05 — 8:25]
Let's now understand how the compiler compiles the template function.
It uses a mechanism called two-phase name lookup.
During the first phase, the compiler is going to check the syntax. It will also check that the
names that do not depend on the template parameter are valid. It will also check that the
operations on the types that do not depend on the template parameter are valid.
During the second phase, the compiler will replace the type it has deduced from the function call
into the function and redo a second round of check to see whether the type exists and that the
operation on them is legal.
Function template
Implicit requirements
Implicit requirements are the sets of operations and properties that a type must support for the
template to compile
template<typename T>
void mySwap(T& a, T& b) {
T tmp = std::move(a);
a = std::move(b);
b = std::move(tmp);
}
What are the implicit requirements for the type T?
🙋♂️🙋♀️
Must be move constructible i.e have a valid T(T&&)
constructor
Must be move assignable i.e have a valid T& operator=(T&&)
assignment operator
💡 C++ 20 formally defines the requirement via the swappable
concept.
💡 Before C++20, document the requirement explicitly using comments.
ℹ️ There are numerous possible implicit requirements
For instance:
T must have a public class member variable T.data
T must have a function member T::fun(int)
T must have a nested type T::TYPE
T must have a static
function member
👉 Implicit requirement is a very important notion
👈
[1:40 — 10:05]
This brings us to one of the most important notions of this talk. Implicit requirements. The
implicit requirements are the set of operations that a type must support for the template to
compile. "Implicit" because they are not stated in the template's signature but are hidden in
the body of the template.
Let's see if you have followed what I said until now. Can anyone tell me what are the implicit
requirements of the `mySwap` template?
T must be move constructible from line 3 and must be move assignable from line 4 and 5.
When implicit requirements aren't met this is the point where the compiler will generate an
error, often with messages that can be difficult to decipher.
Fortunately from C++ 20, we can use concepts to explicitly specify these requirements. I invite
you to follow the talk of Amir Kirsh on Friday.
If you are not using C++20, I seriously recommend you to write the implicit requirements as a
comment in the code.
It is to be noted that there are plenty of different implicit requirements, for instance...
Template argument deduction
General
Templates argument are deduced from the function arguments
For instance
template<typename T>
T myMax(T a, T b) {
return a > b ? a : b;
}
void fun() {
int a = 5, b = 4;
auto c = myMax(a, b);
static_assert(std::is_same_v<decltype(c), int>);
}
Calling myMax(a, b) deduces T = int
[0:50 — 10:55]
Let's move on now to the next topic which is the template argument deduction.
For this section I'm using a simple `myMax` template as an example. It takes two values of the
same type and returns the greater one.
The way argument deduction works is that the compiler looks at the types of the function
arguments and tries to match them to the template parameters.
[next fragment]
In this case, the function arguments `a` and `b` are both of type `int`, so the compiler deduces
that `T` must be `int`.
Template argument deduction
Type deduction
Another example
template<typename T>
T myMax(T a, T b) {
return a > b ? a : b;
}
void fun() {
auto c = myMax(5.1, 5.2);
}
What does T deduced to ?
🙋♂️🙋♀️
[0:45 — 11:40]
Another quick one, keep the pace up. `myMax(5.1, 5.2)` — what is `T`?
[Short pause. Someone will say float. Let them.]
(fragment) `double`. An unsuffixed floating-point literal in C++ is a `double`, never a `float`.
You would need `5.1f`.
Small thing, but it matters on embedded targets where `double` arithmetic may be emulated in
software and cost you a hundred cycles. Deduction gives you exactly what you wrote — no more, no
less. If you wrote the wrong literal, you get the wrong type.
Template argument deduction
Explicit template arguments
Provide the template argument explicitly to override deduction
template<typename T>
T myMax(T a, T b) {
return a > b ? a : b;
}
void fun() {
auto c = myMax<float>(5.1, 5.2);
}
⚠️ Template arguments are never deduced from the return type ⚠️
template<typename T>
T parse(const std::string& text);
int i = parse("42"); // ❌ Does not compile
auto i = parse<int>("42"); // 👍
[1:05 — 12:45]
If you don't want what deduction gives you, you can override it: write the template argument
explicitly. `myMax<float>(5.1, 5.2)`. Now `T` is `float`, deduction is skipped entirely,
and
the `double` literals are converted at the call.
Explicit arguments always win. Deduction only fills in what you did not specify.
(fragment) Now a rule that catches everybody at least once: template arguments are *never*
deduced
from the return type. Look at `parse`. Writing `int i = parse("42")` does not compile — the
compiler
does not look at what you are assigning to. Deduction works left to right, from the arguments
in.
You must write `parse<int>("42")`.
And remember this, because it comes back in the second half: some templates can *only*
be called with explicit arguments.
Template argument deduction
How cv qualifiers are deduced
One question decides everything: is the parameter a reference or not?
template<typename T>
void funcValue(T a);
template<typename T>
void funcRefParam(T& a);
int x = 5;
const int cx = 5;
funcValue(x); // T is deduced as int
funcValue(cx); // T is deduced as int (const is dropped)
funcRefParam(x); // T is deduced as int
funcRefParam(cx); // T is deduced as const int (const is kept)
⚠️ A const T& parameter drops it again
template<typename T>
void funcConstRefParam(const T& a);
funcConstRefParam(cx); // T is int, but a is still const int&
💡 Rule of thumbs 💡
By value → cv-qualifiers are dropped
By reference → cv-qualifiers are preserved
[0:45 — 13:30]
Let's talk about how cv qualifiers are deduced in template argument deduction.
On the left handside I'm defining two function templates: one taking its parameter by value, and
one taking its parameter by reference. How are the cv qualifier deduced?
The answer lies whether the parameter is taken by value or by reference. If it is by value,
cv-qualifiers are dropped. If it is by reference, cv-qualifiers are preserved.
Template argument deduction
Arrays and pointers
template<typename T>
void funcValue(T arg);
template<typename T>
void funcRefParam(T& arg);
int arr[5];
int* ptrArr = arr;
// the pointer remains a pointer
funcValue(ptrArr); // T is deduced as int*
// the reference keeps the int* type
funcRefParam(ptrArr); // T is deduced as int*. But arg is now int*& (dangerous)
// the array decays to a pointer, the size is lost
funcValue(arr); // T is deduced as int*
// the reference keeps the array type
funcRefParam(arr); // T is deduced as int[5]
💡 A reference parameter can even deduce the size as an NTTP 💡
template<typename T, std::size_t N>
constexpr std::size_t arraySize(T(&)[N]) { return N; }
int arr[5];
static_assert(arraySize(arr) == 5);
[1:10 — 14:40]
Let's see now how template argument deduction works with arrays and pointers.
Again we are using the same example function that takes its parameter either by value or by
reference.
On the right side we are defining an array and a pointer to it, and then observing how the
template argument deduction behaves for each case.
let's start first with the pointer
Whenever we pass the pointer by value, it remains a pointer. If we pass it by reference, the
reference keeps the pointer type.
If we pass the array by value, it decays to a pointer and the size information is lost. If we
pass it by reference, the array type is preserved and the size remain.
The fact that the size remain is very beneficial because we can for instance deduce the size as
an NTTP
Template argument deduction
⚠️ Why T& is dangerous with pointers ⚠️
A reference parameter will happily deduce T as a pointer type.
template<typename T>
void dangerous(T& retry) {
// ... some logic
++retry;
}
int retries = 0;
int* retriesPtr = &retries;
// 😱 Forgot the '*' 😱
dangerous(retriesPtr); // T is int*, so retry is int*&
// retriesPtr is now pointing to an invalid memory address
printf("%d\n", *retriesPtr); // 💥 disaster 💥
The pointer is incremented, not retries
⚠️ Think twice before using a non-const reference parameter
💡 Validate the deduced type with static_assert and
<type_traits> , e.g. std::is_pointer_v,
std::is_const_v
💡 From C++20, use concepts
[1:10 — 15:50]
Let's now talk about the danger of reference parameters deducing pointer types.
In this example, we have a function named dangerous that increment its input parameter. Nothing
looks wrong, right?
Let's now define a pointer to an integer and pass it to the dangerous function. Notice that we
actually forgot to dereference the pointer.
This will have the nasty effect that the pointer itself is incremented, not the value it points
to, leading to an invalid memory access. There will not be any compiler error to save you from
this
So think twice before using a non-const reference parameter. Since we don't know what will be
passed to us, it can lead to unexpected and dangerous behavior.
When the deduced type surprises you
A generic myMax
template<typename T>
T myMax(T a, T b) {
return a > b ? a : b;
}
What does the following call return? 🤔
myMax("hello", "world");
🙋♂️🙋♀️
We don't know the result 🤷♂️. Pointer comparison
The instantiated function is
const char* myMax(const char* a, const char* b) {
return a > b ? a : b;
}
[0:40 — 16:30]
Let's now come back on the myMax example and see what is happening with string literals.
Can anyone tell me what the following call will return? 🤔
we actually don't know the result
The best way to understand this is to see what will be the instantiated function by the compiler.
By looking at it we see that a and b a pointer to const char and so we are comparing memory
addresses.
When the deduced type surprises you
Fix: Provide a function overload
A non-templated overload function gives a custom implementation for a specific type.
const char* myMax(const char* a, const char* b) {
return std::strcmp(a, b) > 0 ? a : b;
}
Now myMax("hello", "world") returns the pointer to "world" 👍
💡 Out of reach for a macro: the pre-processor cannot overload 💡
[0:35 — 17:05]
However, fixing this problem isn't difficult
We just need to provide a non-templated function overload for the specific type. And the compiler
is going to pick it.
The fact that we can provide an overload is a clear avantages over macros because we have the
possibility to customize the implementation for a specific type.
Putting templates to work
Generic programming
Iterate over any kind of iterables
template<typename Iterable>
void print(Iterable& c) {
for (const auto& element : c) {
std::cout << element << "\n";
}
}
Works with
Any standard containers (vector, list, set, map, ...)
std::vector<int> v = {1, 2, 3};
print(v);
C++ 20's span and ranges
std::span<int> s = {v};
print(s);
std::vector<int> v = {1, 2, 3, 4, 5};
auto even = v | std::views::filter([](int n) { return n % 2 == 0; });
print(even);
C-style arrays
int arr[] = {1, 2, 3};
print(arr);
Custom types
struct Graph { ... };
// Make Graph printable by adding an overload for print
void print(Graph& g) { ... }
Graph g;
print(g); // 👍
Extensible without ever touching print 👍
[0:50 — 17:55]
Let's now write a generic function.
This print function is very powerful because it can handle any iterable supporting range based
for loops. Thus allowing generic programming.
It works with a vector, a std::span, a view, and even with a C-style array.
This function is also extensible because we can easily add support to a new type which doesn't
support range based for loop to make it printable by simply adding an overload for the `print`
function.
Putting templates to work
"Polymorphism" without inheritance
Polymorphic behaviour between unrelated types
template<typename T>
constexpr int legs(T& ll) { return ll.legs(); }
struct Cow { // not in any class hierarchy
constexpr int legs() { return 4; }
};
struct Duck { // not in any class hierarchy
constexpr int legs() { return 2; }
};
int main()
{
Cow c;
Duck d;
// resolved at compile time, no vtable, no common base
static_assert(legs(c) == 4);
static_assert(legs(d) == 2);
}
Example from Bjarne Stroustrup (apologies for the artificial example)
[1:10 — 19:05]
Our second example is that templates can enable a sort of polymorphic behavior without requiring
inheritance.
It is an example that I read from a post of Bjane Stroustrup. But when I asked him the
permission to use it, he said I need to mention his apology for this example because of the
artificial nature of the example.
In practice, we create a function which calls a member function on any type that provides it,
without requiring a common base class. This is an implicit requirement.
Now let's come back on the example.
We define a templated function legs that calls the legs member function
on any type that provides it. This is the implicit requirements.
Now let's define 2 structs which implement the implicit requirements.
Finally we can see that the templated legs function calls the correct legs member function. The
resolution happened at compile time.
Variadic templates
A variable number of template parameters
The ellipsis ... says which parameter is a pack , and where it
expands.
template<typename T, typename ...Ts>
T sum_all(Ts... args) {
std::array<T, sizeof...(Ts)> arr = { args... }; // pack expansion fills the array
return std::accumulate(arr.begin(), arr.end(), 0);
}
// Any number of arguments can be given to sum_all
auto total = sum_all<int>(1, 2); // total == 3
💡 T doesn't appears in any function parameter → it can only be given
explicitly 💡
[1:45 — 20:50]
So far every template took a fixed number of parameters. Variadic templates lift that to a
variable number.
The whole feature is one token: the ellipsis, three dots. It has exactly two jobs, and once you
see them you can read any variadic code.
But before we dive into the detail, let's have a look at our next example.
The responsibility of the sum_all function will be to sum all the arguments passed to it. This
function takes 2 template parameters: the type of the accumulator (`T`) and a variadic pack of
argument types (`Ts...`).
Now let's talk about the job of the ellipsis.
Job one — to the *left* of a name, in a declaration: "this is a pack". `typename ...Ts` declares
a
pack of types. `Ts... args` declares a pack of function arguments.
Job two — to the *right* of an expression: "expand the pack here".
We see it expanded in the function arguments as well as in the initialization of the std::array.
Declare on the left, expand on the right. That is genuinely the whole rule.
Regarding its call, we need to explicitly specify what T is. That's because T isn't part of the
function parameters and so it must be given explicitly.
Variadic templates
sum_all under the hood 🔎
What the pack expansion really generates
cppinsights.io/s/b7e622a8
Source:
template<typename T, typename ...Ts>
T sum_all(Ts... args) {
std::array<T, sizeof...(Ts)> arr = {args...};
return std::accumulate(arr.begin(), arr.end(), 0);
}
auto total = sum_all<int>(1, 2);
Insight:
template<>
int sum_all<int, int, int>(int __args0, int __args1) {
std::array<int, 2> arr = {{__args0, __args1}};
return std::accumulate(arr.begin(), arr.end(), 0);
}
auto total = sum_all<int>(1, 2);
https://cppinsights.io/s/b7e622a8
⚠️ If sum_all is called with 3 arguments a new version is generated
[1:00 — 21:50]
Back to C++ Insights, because "pack expansion" sounds magical until you see it.
There is no magic. The compiler generated a completely ordinary function taking two `int`s by
value, and the pack expansion became a plain braced initialiser list: `{arg1, arg2}`.
The dots are gone. It is code you could have written by hand.
That is the mental model I want you to keep for the rest of the talk: a pack expansion is
compile-time copy-paste, done by the compiler, with full type checking.
And call it out: if I pass five arguments instead of four, a *second* function is generated. Same
trade-off as before — one instantiation per distinct argument list. But the optimizer helps.
Variadic templates
C++17's fold expressions
Syntax
(pattern op ...);
A fold repeats one expression for each element of the pack.
template<typename ...Args>
void print(Args... args) {
((std::cout << args << " ") , ...);
}
Calling print(1, 2.5, "hello") prints 1 2.5 hello
ℹ️ Before C++17 you peeled the pack off one argument at a time, with a recursive
overload and a non-template base case
[1:00 — 22:50]
Packing into an array works, but usually you want to *do something* with each element. C++17 gave
us fold expressions.
The syntax works as follows. In between mandatory parentheses, we place the pattern we want to
repeat followed by an operator followed by the ellipsis.
A fold repeats one expression for each element of the pack.
I am using the comma operator here, which is the "just do this for each element" fold. Read it
as:
stream out `args`, then a space; repeat for every element.
(fragment) `print(1, 2.5, "hello")` prints `1 2.5 hello`. Three different types, one function, no
recursion.
Before C++17, the technique consists of using recursion.
Variadic templates
Under the hood: one call per pack element 🔎
cppinsights.io/s/16760da2
Source:
template<typename ...Args>
void print(Args... args)
{
((std::cout << args << " ") , ...);
}
print(1, 2.5, "hello");
Insight:
template<>
void print<int, double, const char *>(int __args0,
double __args1,
const char * __args2)
{
(std::cout << __args0 << " "),
(std::cout << __args1 << " "),
(std::cout << __args2 << " ");
}
print(1, 2.5, "hello");
https://cppinsights.io/s/16760da2
https://en.cppreference.com/cpp/language/fold
[0:35 — 23:25]
One more look under the hood, quickly — same lesson, so don't linger.
The fold expanded into three separate, ordinary stream calls, one per pack element, each with the
correct type already resolved. No loop, no recursion, no run-time dispatch. Straight-line code
the
optimiser can chew through.
The cppreference link at the bottom lists all four fold forms — unary/binary, left/right.
Bookmark
it; nobody remembers which is which, including me.
Perfect forwarding
Passing arguments on to another function while preserving their value category (lvalue or rvalue).
template<typename Args>
void relay(Args&& args) {
consume(std::forward<Args>(args));
}
Textbook example: std::vector::emplace_back
template<typename ...Args>
reference emplace_back(Args&&... args);
push_back builds a temporary, then moves it
users.push_back(User{2, "Bob"});
emplace_back calls the constructor in place
users.emplace_back(2, "Bob");
https://godbolt.org/z/334GsGa4n
[1:45 — 25:10]
One more building block we need to go over is perfect forwarding.
The problem: you have a wrapper that takes arguments and passes them to another function. If you
do it naively you lose information — specifically the *value category*. A temporary that arrived
as an rvalue, ready to be moved from, becomes a named lvalue inside your function and gets
copied instead.
You silently turned a move into a copy.
`std::forward` fixes that. Take the parameter as `T&&` where `T` is deduced, and forward
it.
If the caller passed an lvalue, `consume` sees an lvalue; if an rvalue, `consume` sees an
rvalue.
Perfectly preserved — hence the name.
(fragment) The textbook example is `emplace_back`. Its signature is a variadic pack of forwarding
references.
Compare the two calls. `push_back(User{2, "Bob"})` constructs a temporary `User`, then moves it
into the vector, then destroys the temporary.
`emplace_back(2, "Bob")` forwards the raw arguments all the
way down and constructs the `User` *directly in the vector's storage*. One object instead of
two. No
move, no destructor call.
Notice that emplace_back takes the constructor arguments as parameters and not a constructed
object like push_back.
The godbolt link below shows that in the push_back case, a temporary is created and then moved
away while in the emplace_back case, only the constructor is called.
Perfect forwarding
The std::forward Catch
std::forward only works on a forwarding reference : a T&&
where T is deduced by this very function .
template <typename T>
void relay(T&& arg) {
consume(std::forward<T>(arg));
}
⚠️ These are not forwarding references ⚠️
void relay(std::vector<int>&& arg); // fixed type
template<typename T>
void relay(const T&& arg); // const rvalue
Two more traps:
⚠️ Using arg again after forwarding it → moved-from bug
⚠️ std::move instead of std::forward → moves lvalues
unconditionally
Forwarding references are also called universal references : Scott
Meyers
[1:55 — 27:05]
Here is the catch.
`std::forward` only does its job on a genuine *forwarding reference*. And a forwarding reference
is
very precisely defined: a `T&&` where `T` is a template parameter deduced *by this very
function*. Both halves matter — I have highlighted them.
It is very confusing because it looks like a plain rvalue reference, but it is actually a
forwarding reference. Just because it is a template parameter.
(fragment) These are *not* forwarding references, even though they have two ampersands. The first
takes a fixed type — that is a plain rvalue reference; it only binds to rvalues, nothing is
deduced.
The second is `const T&&` — the `const` disqualifies it. The pattern has to be exactly
`T&&`.
(fragment) Two more traps. One: after you forward an argument, treat it as moved-from. Using it
again is a bug — and a quiet one, because a moved-from object is valid, just unspecified. Two:
don't
reach for `std::move` instead of `std::forward` here. `move` casts unconditionally, so it will
steal
from a caller's lvalue that they still intend to use. `forward` moves only when the caller gave
you
an rvalue.
Rule of thumb: `std::move` on a concrete rvalue reference, `std::forward` on a forwarding
reference.
Never cross them.
(fragment) Terminology note: Scott Meyers coined "universal reference" before the standard
settled
on "forwarding reference". Same thing, you will see both.
Class/struct template
Defines a family of classes
A class template describes one class blueprint that the compiler instantiates for each type.
template<typename T>
class Stack {
private:
std::vector<T> data;
public:
void pop() { data.pop_back(); }
const T& top() const { return data.back(); }
bool empty() const { return data.empty(); }
std::size_t size() const { return data.size(); }
template <typename ...U>
void emplace(U&&... args) { data.emplace_back(std::forward<U>(args)...); }
};
Stack<int> intStack;
Stack<float> floatStack;
Stack<int> and Stack<float>
are different types generated from the same template.
Because Stack<int> and Stack<float> are types,
they can be given as template parameter values
[1:40 — 28:45]
Last building block: class templates. Everything you know about function templates carries over.
We are building a stack abstraction using a vector as an inner data structure.
Let's go through the code step by step.
(lines 1-2) We tell the compiler we are building a templated class by using the keyword
`template<typename T>` before the class definition.
(lines 3-4) `T` is usable anywhere inside — here as the element type of the underlying vector.
(lines 5-12) Ordinary member functions. T can be used anywhere as well.
Notice also the function emplace which uses perfect forwarding to emplace the element in place
into the vector.
(lines 15-16) Instantiate with `int` and with `float`.
(fragment) And here is the thing to internalise: `Stack<int>` and `Stack<float>` are
different, unrelated *types*. Not two instances of one type — two types. You cannot assign one
to
the other. They have no common base class.
(fragment) And because they are types, they can themselves be used as template arguments. That is
how you get `std::vector<std::vector<int>>` — and it is the mechanism behind CRTP,
which
is a few slides away. Types in, types out. Templates compose.
Perfect forwarding
The std::forward Catch reloaded
⚠️ Not all T&& are forwarding references
template <typename T> // T deduced HERE, at instantiation
class Stack {
// ... other member functions ...
// args looks like a forwarding reference but T is already fixed
void emplace(T&&... args) { data.emplace_back(std::forward<T>(args)...); }
};
Example of instantiation: Stack<int> s;
struct Stack {
// ... other member functions ...
// ⚠️ Only accept rvalues references of type int
void emplace(int&&... args) { data.emplace_back(std::forward<int>(args)...); }
};
Fix: make emplace
templated
template <typename T>
struct Stack {
// ... other member functions ...
template <typename ...U>
void emplace(U&&... args) {
consume(std::forward<U>(args)...);
}
};
[1:20 — 30:05]
You might wonder why `emplace` is templated and the answer lies in the forwarding reference
rules.
Look at `emplace`. It takes `T&&`. There is a `typename T` right above it. It looks
exactly
like a forwarding reference. It is not.
Why? Because `T` is deduced at *class* instantiation — when you write `Stack<int>` — not by
the member function. By the time anyone calls `emplace`, `T` is already fixed. And the rule
says:
deduced *by this very function*.
(fragment) To convince ourselves, let's look at what the compiler actually generates for
`Stack<int>`. There it is in black
and white: `void emplace(int&&...)`. The function only accepts r-value references.
The bug is there but your code will compile.
(fragment — the fix box) The fix is one extra line: give the member function *its own* template
parameter, `U`. Now `U` is deduced at the call, `U&&` is a real forwarding reference,
and
`std::forward<U>` behaves.
Class/struct template
Function member instantiation
void useStack() {
Stack<int> s{};
s.emplace(5);
std::printf("%d\n", s.top());
}
https://godbolt.org/z/n7K1PKdrs
Only the called member functions are instantiated
⚠️ A member function that would not compile for T stays silent until
someone calls it
[0:40 — 30:45]
One more property of class templates, and it surprises people.
Looking at the example, we create a stack of int and we call the functions emplace and top.
Only the code of the called function is actually instantiated.
If you remember the two-phase initialization mechanism, this can hurt you.
It can be that your code compiles successfully for a long time, and then suddenly fails when an
unused member function is called.
Class/struct template
Curiously Recurring Template Pattern (CRTP)
A class can pass itself as the template argument when deriving from it.
ℹ️ The base class knows the type CHILD of the class deriving from it
template <typename CHILD>
struct Base {
void greet() {
static_cast<CHILD*>(this)->greet_impl(); // Base knows the derived type
}
};
struct Derived : public Base<Derived> { // pass yourself as the argument
void greet_impl() { std::puts("hello"); }
};
Derived d;
d.greet(); // prints "hello", no virtual call
💡 Static polymorphism, resolved at compile
time 💡
[1:50 — 32:35]
The last technique of today is the Curiously Recurring Template Pattern (CRTP).
This allows a parent class to know the exact static type of the derived class.
Let's look at an example.
(lines 1-2) `Base` looks like a normal template class, but it takes a type parameter I have named
`CHILD`.
(lines 3-5) Because we know that CHILD is the type of the derived class, we can safely cast
`this` to `CHILD*` and call member functions on it.
(lines 8-10) Now let's define the derived class. It inherits from the base class and passes
itself as the template argument.
The derived class also implements the `greet_impl` function that the base class will call.
(lines 12-13) Let's now create an instance of the Derived class and call the greet function of
the base class. Even though we call the base class function, a function of the derived class is
invoked.
(fragment) This is static polymorphism with an inheritance *shape*: the base provides shared
behaviour, the derived class provides the specifics, and it is all resolved at compile time.
(fragment) Say the key line again: the base class knows the type of the class deriving from it.
That one fact is what makes the third example — the hardware register abstraction — possible.
Keep
it warm; we use it in about ten minutes.
Let's see templates in action
🚀🚀🚀🚀🚀🚀🚀
[0:20 — 32:55]
We now have all the tools in our toolbox to have a look at our 3 examples so let's get started.
The 3 examples will be structured as follows.
Generic serializer using a policy-based design
Design a dump_object function to serialize an object into
different formats (e.g., JSON, YAML)
It should use policy classes (sinks) to determine the output format
No inheritance required
Easy to extend with new sinks without modifying the function itself
[0:50 — 33:45]
Example one: a generic serializer using policy-based design.
We want a `dump_object` function that serializes an object into different formats — JSON, YAML,
maybe a binary protocol later.
The output format is chosen by a *policy class* — a sink.
No inheritance. No abstract base `Serializer`, no virtual `write_field`.
And adding a new format must not require touching `dump_object`.
Generic serializer using a policy-based design
template <typename Sink>
void dump_object(const Object& d)
{
Sink::begin_object("object");
Sink::field("id", d.id());
Sink::field("name", d.name());
Sink::end_object();
}
struct JsonSink {
static void begin_object(std::string_view n) {
std::printf("\"%s\": {\n", n.data());
}
static void field(std::string_view k, int v) {
std::printf(" \"%s\": %d,\n", k.data(), v);
}
static void field(std::string_view k, std::string_view v) {
std::printf(" \"%s\": \"%s\",\n", k.data(), v.data());
}
static void end_object() { std::printf("}\n"); }
};
struct YamlSink {
static void begin_object(std::string_view n) {
std::printf("%s:\n", n.data());
}
static void field(std::string_view k, int v) {
std::printf(" %s: %d\n", k.data(), v);
}
static void field(std::string_view k, std::string_view v) {
std::printf(" %s: %s\n", k.data(), v.data());
}
static void end_object() {}
};
Main function
int main() {
Object apple{1, "apple"};
std::printf("---------JSON-----------\n");
dump_object<JsonSink>(apple);
std::printf("---------YAML-----------\n");
dump_object<YamlSink>(apple);
}
https://godbolt.org/z/crqP67bTW
Output:
---------JSON-----------
"object": {
"id": 1,
"name": "apple",
}
---------YAML-----------
object:
id: 1
name: "apple"
[1:30 — 35:15]
Here is the whole thing. Eight lines at the top.
`dump_object` takes a `Sink` as a template parameter and calls static functions on it:
`begin_object`, `field`, `end_object`. That is all. `Sink` is a *type*, not an object — there is
no
parameter, nothing is passed at run time.
The implicit requirements are: a static `begin_object`, a static `field` for each value type, a
static `end_object`.
(fragment — JsonSink) Here is a sink that satisfies them, emitting JSON. Plain struct, static
members, no base class, no virtual.
(fragment — YamlSink) And here is YAML. Note `end_object` is *empty* — YAML has no closing brace.
With a virtual interface you would still pay for that call. Here it inlines to literally
nothing.
(fragment — main + output) And main: same object, two formats, two lines. The output is on the
right. Godbolt link if you want to poke at it.
In this case the decision happened at compile time and we didn't need to rely on any vtables.
Generic abstraction layer
Register bank manager
Register bank is used to program the hardware
2 set of registers bank exists: Active and Standby
Driver writes the standby bank using write(...)
write takes an unspecified number of parameters
In our case it is write(name, value)
Driver calls commit() to swap the active and standby banks
Bank-swap policy is decoupled from register access logic
The driver code is unaware of the active bank
[1:15 — 36:30]
Example two, straight out of embedded work: a register bank manager.
When programming the HW, we typically have to deal with active and standby banks of configuration
registers. This allows the driver to write to the standby bank without affecting the currently
active configuration.
We want to design a write function that the driver will use to update the standby bank. The
function should be flexible enough to handle different register layouts for various peripherals
and thus take an unspecified number of parameters.
We also want a commit function that swaps the active and standby banks atomically.
The bank-swap policy is decoupled from the register access.
And the driver is unaware of which bank is currently active.
Generic abstraction layer
Register bank manager
Driver code
int main()
{
ActiveStandbyBank<MotorConfig> mgr(MotorConfig{});
// Write to the current standby bank i.e bank 0
if(!mgr.write("speed_rpm", 1500)) {
std::cerr << "Cannot write speed_rpm\n";
}
if(!mgr.write("torque_limit", 80)) {
std::cerr << "Cannot write torque_limit\n";
}
mgr.commit();
// Now writing to bank 1 (the new standby bank)
if(!mgr.write("speed_rpm", 1000)) {
std::cerr << "Cannot write speed_rpm\n";
}
}
Register access logic implementation
class MotorConfig
{
private:
struct Config { uint16_t speed_rpm = 0;
uint16_t torque_limit = 0; };
Config m_banks[2]{};
public:
bool write(uint8_t bank_id, std::string_view field, uint16_t value)
{
if (field == "speed_rpm") {
m_banks[bank_id].speed_rpm = value;
return true;
} else if (field == "torque_limit") {
m_banks[bank_id].torque_limit = value;
return true;
}
return false;
}
void commit(uint8_t bank_id)
{
// action to commit the bank
}
};
[1:50 — 38:20]
Let's now have a look at how the driver code interacts with the register bank manager.
(line 3) Create an `ActiveStandbyBank` wrapping a `MotorConfig`. (lines 5-11)
Write two fields by name. Look at what is *missing*: no bank index anywhere. (line 13) Commit.
(lines 15-18) Write again — and this now goes to the other bank, automatically. The client never
knew, and cannot get it wrong.
Also note the return value is propagated — `write` returns a `bool` and we check it. Hold that
thought, it matters in a second.
HAL implementation, bottom. (line 1) A completely plain class. No base class, no virtual, no
`#include` of the manager. It does not know `ActiveStandbyBank` exists.
(lines 3-6) Its own `Config` struct and an array of two banks. (line 8) And here is the contract:
`write` takes a `bank_id` as its *first* parameter, then whatever else this particular
peripheral
needs — here a field name and a value. (lines 9-18) The body is boring hardware code. (lines
20-23)
And `commit` likewise takes a `bank_id` first.
So the entire implicit requirement is: "provide `write` and `commit` whose first parameter is a
bank
id". Everything after that is yours. A different peripheral could take a register offset and a
32-bit value, or three arguments, or none.
That flexibility is exactly what an abstract base class cannot give you — a virtual function has
one
fixed signature.
Generic abstraction layer
Register bank manager
ActiveStandbyBank implementation
template <typename REGISTER_ACCESS_LOGIC>
class ActiveStandbyBank
{
private:
REGISTER_ACCESS_LOGIC m_register_access_logic{};
uint8_t m_active_id = 0;
public:
explicit ActiveStandbyBank(REGISTER_ACCESS_LOGIC register_access_logic)
: m_register_access_logic(std::move(register_access_logic)){}
uint8_t active_id() const { return m_active_id; }
uint8_t standby_id() const { return (m_active_id + 1) % 2; }
template <typename ...Args>
bool write(Args&&... args)
{
return m_register_access_logic.write(standby_id(), std::forward<Args>(args)...);
}
template <typename ...Args>
void commit(Args&&... args)
{
const auto current_standby_id = standby_id();
m_active_id = current_standby_id;
m_register_access_logic.commit(current_standby_id, std::forward<Args>(args)...);
}
};
https://godbolt.org/z/zd5EG59P4
[1:15 — 39:35]
Now let's have a look at the ActiveStandbyBank implementation which is the star of the show.
It is 27 lines and it works with ANY register logic that satisfies the expected interface.
(lines 1-2) A class template over `REGISTER_ACCESS_LOGIC`.
(lines 4-6) It owns the register logic by value — no pointer, no indirection, and the compiler
can see right
through it — plus one byte of state: which bank is active.
(lines 8-12) Constructor, and the trivial bank arithmetic. This is the policy, written once.
(lines 14-18) `write`. This is the interesting part.
We got our forwarding reference pack.
Then on line 17, we call `m_register_access_logic.write` with the standby id and perfectly
forward the arguments.
(lines 20-26) `commit` is the same shape.
It grabs the arguments using a forwarding reference pack.
It is to be noted that in our case the commit function took no argument so the parameter pack
was empty. But that is ok.
Flexible and safe register accesses
Concept
Hardware is programmed through registers
Each register has fields of different sizes
Modifying a field is a generic read-modify-write (RMW) algorithm Read
the register, clear the field's bits, set the new value, write it back
For instance: ARM PL011 UART has a register called UARTLCR_H
Bits
Name
15:8
Reserved
7
SPS - Stick parity select
6:5
WLEN - Word length
4
FEN - Enable FIFOs
Bits
Name
3
STP2 - Two stop bits select
2
EPS - Even parity select
1
PEN - Parity enable
0
BRK - Send break
[0:50 — 40:25]
And finally the last example which is my favorite.
We have talked about registers and we know they are used to program the hardware.
In order to save memory space, several settings are grouped into a single register using fields.
They have a different size and position within the register.
To update the fields we need to use a read-modify-write operation: read the current value of the
register, modify the specific field, and write the new value back.
For instance here we see the fields of a UARTLCR_H register from the ARM PL011 UART.
Flexible and safe register accesses
Goal
Provide a Register::write(...) function to write the fields
Multiple fields can be written at once i.e multiple parameters
We can choose how many field to write at once
The order of parameters doesn't matter
Type safe
Client code
UARTLCR_H lcrh{registerAddress};
lcrh.write(
// 8 data bits
UARTLCR_H::WLEN{0b11},
// enable FIFOs
UARTLCR_H::FEN{0b1},
// no parity
UARTLCR_H::PEN{0b0}
);
UARTLCR_H lcrh{registerAddress};
lcrh.write(
// enable FIFOs
UARTLCR_H::FEN{0b1},
// no parity
UARTLCR_H::PEN{0b0},
// 8 data bits
UARTLCR_H::WLEN{0b11}
);
[1:35 — 42:00]
Here is what I want the API to look like. Requirements first, then the client code.
One `write` function. It must accept multiple fields at once — one read-modify-write cycle for
the
whole update, which for hardware is not just faster, it is *correct*: the register is never in a
half-configured state. The order of the fields must not matter — they are named, not positional.
And
it must be type safe.
(fragment — left) So the client writes this. Each field is a small named object carrying its
value:
`WLEN{0b11}` for 8 data bits, `FEN{1}` to enable FIFOs, `PEN{0}` for no parity. It reads like
the
datasheet. No shifts, no masks, no hex constant you have to trust.
(fragment — right) And the same three fields in a different order does exactly the same thing.
Now — the type safety claim. What I want is for `UARTLCR_H::WLEN` and, say, `UARTCR::WLEN` to be
different types, so that passing the wrong one is a *compile error*, not a silent corruption of
a
neighbouring field. That is the day-two-with-an-oscilloscope bug, and I want the compiler to
catch
it.
Let's build it. Two slides.
Flexible and safe register accesses
UARTLCR_H register definition
// Use CRTP to give the parent class the type of the register it abstracts
struct UARTLCR_H : public Register<UARTLCR_H> {
UARTLCR_H(volatile uint32_t* regAddr):Register{regAddr} {}
struct WLEN {
using REG = UARTLCR_H;
static constexpr uint32_t FIELD_OFFSET = 5;
static constexpr uint32_t FIELD_SIZE = 2;
static constexpr uint32_t FIELD_MASK = ((1 << FIELD_SIZE) - 1) << FIELD_OFFSET;
uint32_t value;
WLEN(uint32_t value):value{value} {}
};
struct PEN {
using REG = UARTLCR_H;
static constexpr uint32_t FIELD_OFFSET = 1;
static constexpr uint32_t FIELD_SIZE = 1;
static constexpr uint32_t FIELD_MASK = ((1u << FIELD_SIZE) - 1) << FIELD_OFFSET;
uint32_t value;
PEN(uint32_t value):value{value} {}
};
// Other fields implementation...
};
ℹ️ Each sub-structure implements the implicit requirements of the register
abstraction
[1:30 — 43:30]
Let's first have a look at the UARTLCR_H register implementation.
(lines 1-2) We are using the CRTP pattern to tell the parent class the type of the register it
abstracts. We will see in a moment how useful that is.
(line 3) Constructor takes the memory-mapped address.
Then we have our first field definition.
Each field of the register will be a nested struct.
The structure will have: a nested type. This type is actually the type of the register that the
fields belong to. Three static constexpr values: the offset, the size, and the mask computed
from them. And finally, a runtime `value` member which is the value that we want to write into
the field.
Then we do the same for all the other fields, for instance PEN.
(fragment) The point: each nested struct satisfies the implicit requirements of the abstraction —
a
nested type `REG`, three static constexpr values, and a `value` member. That is the contract.
And
notice it includes a *nested type* — remember I flagged that back on the implicit requirements
slide. This is where it pays off.
Flexible and safe register accesses
Register abstraction implementation
template<typename CHILD>
class Register{
private:
volatile uint32_t* value; // Points to the actual memory-mapped register
public:
Register(volatile uint32_t* regAddr):value{regAddr}
{}
uint32_t& operator*() {
return *value;
}
template <typename ...FIELD>
void write(FIELD... field)
{
static_assert((std::is_same_v<typename FIELD::REG, CHILD> && ...),
"Using field from the wrong register");
uint32_t curVal = *value; // Read value
auto applyField = [&](auto f)
{
using FIELD_TYPE = decltype(f);
curVal &= ~FIELD_TYPE::FIELD_MASK; // Clear the field
curVal |= (f.value << FIELD_TYPE::FIELD_OFFSET) & FIELD_TYPE::FIELD_MASK; // Set the field's value
};
(applyField(field), ...);
*value = curVal; // Write back value
}
};
👍 Any number of fields
👍 Fields order doesn't matter
💡
static_assert + <type_traits> ensure fields from the right
register are used 💡
ℹ️ The fold expression executes the lambda for each field
[1:50 — 45:20]
Here comes the register abstraction. Let's go into it bit by bit.
(lines 1-2) First, we receive the child register type thanks to CRTP.
(lines 3-7) We then define some member variables, including a `volatile` pointer to the
memory-mapped register.
We also take care of defining an operator* that will allow convenient access to the underlying
register value.
Let's now have a look at the write function.
(lines 8-9) `write` takes a variadic pack of fields. That is how you get "any number, any order".
(lines 11-12) We enable type safety using static_assert and std::is_same_v. std::is_same_v checks
whether the two types given as template arguments are the same.
Thanks to the nested type REG of the field, and thanks to the CHILD template parameter, we can
check at compile time that the field belongs to the correct register.
(lines 14-19) Finally we use read-modify-write to clear the field value and set it to the new
one.
In order to apply the algorithm to each value, we make use of a lambda from line 20 to 25.
From the field type we can apply the right binary operation to achieve what we need.
And finally on line 26, we use the fold expression to call the lambda on each field.
Flexible and safe register accesses
What does all of this cost?
The abstraction executes a lambda for each field...
uint32_t testValueAbstraction() {
uint32_t v{0};
UARTLCR_H lcrh{&v};
lcrh.write(
UARTLCR_H::WLEN{0b11},
UARTLCR_H::FEN{0b1},
UARTLCR_H::PEN{0b0}
);
return *lcrh;
}
uint32_t testValueDirect() {
volatile uint32_t v{0};
v = (v & ~0x72) | 0x70;
return v;
}
Is the abstraction slower, faster or the same?
🙋♂️🙋♀️
[0:30 — 45:50]
Now you might wonder. We've got one lambda call per field because of the fold expression. What is
the cost of all of this?
Let's compare it!
I put on the left the abstraction, and on the right the hand-written direct register access with
hardcoded value.
Last quiz of the day.
Flexible and safe register accesses
True zero cost abstraction
uint32_t testValueAbstraction() {
uint32_t v{0};
UARTLCR_H lcrh{&v};
lcrh.write(
UARTLCR_H::WLEN{0b11},
UARTLCR_H::FEN{0b1},
UARTLCR_H::PEN{0b0}
);
return *lcrh;
}
testValueAbstraction():
sub sp, sp, #16
str wzr, [sp, 12]
ldr w0, [sp, 12]
and w0, w0, -3
orr w0, w0, 112
str w0, [sp, 12]
ldr w0, [sp, 12]
add sp, sp, 16
ret
uint32_t testValueDirect() {
volatile uint32_t v{0};
v = (v & ~0x72) | 0x70;
return v;
}
testValueDirect():
sub sp, sp, #16
str wzr, [sp, 12]
ldr w0, [sp, 12]
and w0, w0, -3
orr w0, w0, 112
str w0, [sp, 12]
ldr w0, [sp, 12]
add sp, sp, 16
ret
https://godbolt.org/z/dfKK4nKoh
[0:20 — 46:10]
Looking at the assembly output, we can see that both the abstraction and the direct register
access result in identical instructions. This demonstrates the zero-cost nature of our
template-based abstraction.
Conclusion: Key Takeaways
🔑 Template is a very powerful feature of C++
🔑 😈 Beware of the implicit requirements 😈
🔑 Replace your macros with templates
🔑 Generic programming
🔑 Allows extensibility
🔑 Good template abstraction can simplify client code without compromise on
performance
Don't be scared of templates 👻
[1:45 — 47:55]
Let's wrap up. Six takeaways.
(1) Templates are a very powerful feature of C++ — and you have now seen that the powerful part
comes from a handful of simple pieces, not from the scary ones.
(2) Beware of implicit requirements. They are the mechanism behind everything we built today,
*and*
the source of the worst error messages. Document them in comments, or state them with concepts
if
you are on C++20.
(3) Replace your macros with templates. The most actionable item on this list — you can do it
this
week.
(4) Generic programming: one implementation, many types, including types that did not exist when
you
wrote it.
(5) Extensibility — non-intrusive extensibility. New sink, new HAL, new register: you add code,
you
never modify the abstraction.
(6) And a good template abstraction simplifies client code without compromising performance. We
proved that one with the assembly.
(fragment) So: don't be scared of templates. Start small — turn one macro into a function
template
this week. That is genuinely how it starts.
The end!
[0:20 — 48:15]
For the record, I gave a small lie at the start of the presentation. While giving the talk to my
cat she decided to take a nap on me...
Slides available at: laurentcarlier.com/b2b-template-cppcon-2026/
[Q&A — 53:00 to 60:00]
KEY POINTS — prepared answers
Compile times? Real cost. Keep templates thin, push type-independent logic into
non-template helpers, explicit instantiation when the type set is fixed. Measure first
When still virtual? Type set not known at compile time: plugins, heterogeneous
containers, ABI boundaries. Compile-time polymorphism complements runtime, it does not
replace it
Code bloat? Often overstated — the linker folds identical instantiations, inlining
usually wins it back. Measure your binary
Concepts instead? Yes if C++20 — they make the implicit requirements explicit and the
errors readable. Nothing today changes
All compilers? Godbolt uses GCC AArch64 -O2; same with Clang and MSVC.
At -O0 of course not — zero cost assumes the optimiser is on
If Q&A dries up: offer to walk back through the register abstraction slide
Thank you. The slides, with every Godbolt and C++ Insights link, are at the address at the
bottom.
Happy to take questions.