Writing Efficient C++ Code

(asawicki.info)

61 points | by ibobev 1 day ago

4 comments

  • hn_submit 1 hour ago
    I write in C++ almost every day but never have the need to optimize for speed. Even when you write straightforward code it's already blazingly fast.
    • gbin 38 minutes ago
      It is probably very domain specific. In robotics for example everything is a zero sum game: CPU, memory bandwidth, GPU, battery life etc ... So it is really a topic, probably true for anything embedded actually. Some other offline applications: HFT, Telco etc.. I wish the GUI apps devs respect more the laptop resources they are running on, don't get me started on the 4 instances of chrome I need to run just for discord, signal etc ...
    • glouwbug 10 minutes ago
      True, but moving from a list of unique polymorphic pointers to a std::variant gains you at least a 2-3x speed up in terms of TLB and cacheline locality. From there, swapping to SOA will net you another 4-8x, so you're looking at nearly 25x improvement by going data first. That may not matter in the unique case of say, games, where rendering a million entities will dwarf the cost of SIMD processing a million entities, but in something like numerical simulations (fluids) or quant it will be warmly welcomed
    • cjbgkagh 40 minutes ago
      I rarely use C++ but when I do it is for speed. It’s not uncommon that carefully crafted intrinsics can 10x the straightforward naive implementation.
  • 112233 1 hour ago
    "This article was originally published in Polish in issue 4/2013" — a lot of excellent advice. Sad to see C++ have moved in last decade in a direction that makes writing efficient, simple low level code harder and harder :(
    • jandrewrogers 1 minute ago
      Writing clear, concise, and efficient code in C++ has never been simpler or easier. The improvements in C++ over the last 15 years have been qualitative.

      So many complex, esoteric, and difficult to maintain incantations that used to be required for efficient code generation are no longer necessary.

    • jll29 57 minutes ago
      I think it has become EASIER: for instance, since C++23 Rust-like move semantics can be used, which provides the compiler with extra information that can be leveraged for the generation of better code.

      Or take constexpr - it permits to move computations to compile time that are complex and in older versions either had to be done at runtime, or an ugly workaround had to be used (e.g. assigning a mysterious literal pre-computed in another run or by hand).

      • creata 37 minutes ago
        > C++23 Rust-like move semantics can be used

        What C++23 feature allows that?

    • fooblaster 1 hour ago
      How? you can write exactly the same low level code today.
      • beached_whale 1 hour ago
        My thought too.

        There are so many things that are expressible in C++ now that could not be without writing much more code or using per-compilation tools back then. The ability to run code at compile time that is not run at runtime is huge, #embed lets us make other tools output available without linker scripts or compiler specific tools that.

        Also, most of the code from the past still works(from 10 years ago definitely works)

      • AlotOfReading 54 minutes ago
        Shot in the dark, but maybe the OP is referring to the fact that these code conventions are explicitly discouraged by the C++ core guidelines. The SoA example falls afoul of the rule requiring T* to be used only for singular object pointers, for example.
        • cjbgkagh 42 minutes ago
          Not a regular C++ programmer but wouldn’t you use std::span here instead? Sure it’ll carry a few redundant lengths but it makes using functions that take spans easier. When I do write C++ it’s usually for speed so I’m often working at the intrinsics level, though AI has gotten good enough at it that I now generally delegate this work to an agent.
          • jandrewrogers 8 minutes ago
            Depending on the specific code, the compiler may even eliminate the redundant lengths.
  • tug2024 1 hour ago
    [dead]
  • MaxBarraclough 41 minutes ago
    There's no mention of branch prediction, or context switching, or synchronisation. Depending on what you're doing, they could be very consequential. There's only very brief mention of parallelisation with threads and with SIMD.

    High-performance programming is a big topic. The scope is far too broad for a single blog post, which naturally gives only cursory discussion of C++ and computer architecture. The article isn't bad considering, but I do think it's the wrong format. A blog series, or even a book, would be more fitting.

    • glouwbug 16 minutes ago
      Learn which instructions SIMD nicely (sqrt / fabs, etc). Use ternaries in loops for masking. Use trig identities and lookup tables (don't recompute sin(3t) when you can use two vector multiples using a table of sin(t) eg. sin(t) * sin(t) * sin(t)). Use divisible constexpr constants in loops to eliminate the SIMD tail. Be careful with type casts and floats. `float x; x += 0.5` will introduce *cvt instructions even if the compiler statically knew better otherwise (use 0.5f). Compile with --fast-math and friends so errno doesn't invalidate your SIMD pipeline.
      • MaxBarraclough 1 minute ago
        That has a similar problem to the article, it's trying to fit far too much into too small a format.

        What you've written makes sense to someone who already has a solid understanding of SIMD and of C++, but the target audience is people who don't. For them, each point needs a much lengthier explanation.

      • creata 4 minutes ago
        Most applications (including most applications that care about numerical performance) should not use -ffast-math.