Computational Performance Archives - Johnny's Software Lab

A story of a very large loop with a long instruction dependency chain

February 29, 2024February 29, 2024Ivica BogosavljevićComputational Performance, Low Level Performance, Performance, VectorizationLeave a Reply

A story of a very large loop with a long instruction dependency chain.

When an instruction depends on the previous instruction depends on the previous instructions… : long instruction dependency chains and performance

September 24, 2022February 29, 2024Ivica BogosavljevićComputational Performance, Low Level Performance, PerformanceLeave a Reply

This post has a second part, the same problem is solved differently. Read more. In this post we investigate long dependency chains: when an instruction depends on the previous instruction depends on the previous instruction… We want to see how long dependency chains lower CPU performance, and we want to measure the effect of interleaving…

Read

Vectorization, dependencies and outer loop vectorization: if you can’t beat them, join them

March 13, 2022August 14, 2022Ivica BogosavljevićComputational Performance, Low Level Performance4 Replies

As I already mentioned in earlier posts, vectorization is the holy grail of software optimizations: if your hot loop is efficiently vectorized, it is pretty much running at fastest possible speed. So, it is definitely a goal worth pursuing, under two assumptions: (1) that your code has a hardware-friendly memory access pattern1 and (2) that…

Read

When vectorization hits the memory wall: investigating the AVX2 memory gather instruction

December 24, 2021March 19, 2022Ivica BogosavljevićComputational Performance, Low Level Performance, Performance4 Replies

For all the engineers who like to tinker with software performance, vectorization is the holy grail: if it vectorizes, this means that it runs faster. Unfortunately, many times this is not the case, and the results of forcing vectorization by any means can mean lower performance. This happens when vectorization hits the memory wall: although…

Read