The basics with teeth: reserve before bulk push_back, emplace_back to construct in place, and noexcept move operations so vector growth moves instead of copies.
v.reserve(n);
for (auto& r : rows) v.emplace_back(r.id, r.name);
Heterogeneous lookup (std::less<>) stops find("literal") from constructing a temporary std::string per call.
std::map<std::string, Val, std::less<>> m;
m.find(std::string_view(key));
False-sharing control gets names: hardware_destructive_interference_size for padding hot atomics apart; string_view/shared_mutex cut copies and writer contention.
struct alignas(std::hardware_destructive_interference_size) Counter {
std::atomic<long> n;
};
[[likely]]/[[unlikely]] annotate cold branches, [[no_unique_address]] compresses layout, <bit> ops map to single instructions, and ranges views fuse passes lazily.
[[assume(expr)]] feeds the optimizer invariants, std::unreachable() prunes impossible paths, flat_map/flat_set trade pointer-chasing trees for cache-friendly sorted vectors.
[[assume(n % 4 == 0)]];
for (int i = 0; i < n; i += 4) step4(i);
std::simd makes data parallelism portable vocabulary, trivial relocation turns container growth into memcpy, and fetch_max/fetch_min shrink CAS loops.
std::simd<float> x(&a[i], std::simd_flags{});
auto y = x * k + b;
y.copy_to(&out[i], std::simd_flags{});
C++29direction · not adopted
Direction: SIMD algorithm integration and relocation-aware library optimizations continue; measure-first remains the only universal idiom.