BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CERN//INDICO//EN
BEGIN:VEVENT
SUMMARY:[Online] Core-Level Performance Engineering
DTSTART:20261005T070000Z
DTEND:20261005T150000Z
DTSTAMP:20261010T223600Z
UID:indico-event-208@indico.ecap.work
DESCRIPTION:Core-Level Performance Engineering\n\nSchedule & Format\n\nDat
 e: 2026\, Oct 5\nTime: 9:00 - 17:00 CE(S)T\nLocation: Online via Zoom\nLan
 guage: English\n\n \nRegistered participants will receive the video confe
 rencing link via email on the day before the course.\nFurther information 
 about the course (content\, learning objectives\, prerequisites\, certific
 ation\, etc.) can be found on the course page.\nInstructor\nDr. Georg Hage
 r\, NHR@FAU and Jan Laukemann\, NHR@FAU.\nThis course is organized by Erla
 ngen National High Performance Computing Center (NHR@FAU).\nCourse Descrip
 tion\nMany developers invest heavily in parallelism while overlooking the 
 efficiency of their serial code. Slow serial code tends to scale well in t
 erms of parallel speedup\, which can mask the fact that compute resources 
 are being wasted. This course closes that gap by conveying a thorough unde
 rstanding of the interactions between software and hardware at the level o
 f a single CPU core and the L1 cache. Topics include superscalar out-of-or
 der execution\, instruction throughput\, critical path and loop-carried de
 pendencies\, and the architectural differences between x86 and ARM process
 ors. Participants also learn to read and interpret compiler-generated asse
 mbly in AT&T (x86) and AArch64 syntax\, and to apply the Open Source Archi
 tecture Code Analyzer (OSACA) together with the Compiler Explorer to asses
 s and model performance properties.\nPrerequisites\nKnowledge\n\nProgrammi
 ng experience in C or C++ at a level sufficient to read and understand sim
 ple loop kernels\nA basic understanding of how CPUs work (registers\, inst
 ruction execution\, data transfers) and what ‘machine instructions’ ar
 e.\nA basic understanding of the Roofline model is recommended. You can fi
 nd some information here (lecture slides) and here (publication by S. Will
 iams).\nSome experience in using the Compiler Explorer is recommended but 
 not required. You can watch a two-part intro by Matt Godbolt: part 1 part 
 2\n\nTechnical\n\nA modern web browser\n\nLearning Outcomes\n\n\nAfter com
 pleting this course\, you will be able to:\n\nhave a good grasp of the out
 -of-order processing capabilities of modern server CPUs\, specifically x86
  and Arm-based server processors such as the Intel Sapphire Rapids and Fuj
 itsu A64FX CPUs\,\nunderstand the concepts of instruction latency\, instru
 ction throughput\, loop body critical path\, and loop-carried dependencies
  and how they impact loop kernel performance\,\nunderstand how SIMD instru
 ctions (i.e.\, vectorized code) can boost performance and how SIMD and pip
 elining/out-of-order processing interact\,\nbe able to work with compiler-
 generated assembly code\,\nbe able to employ Godbolt’s Compiler Explorer
  and the OSACA tool for loop-level in-core performance analysis and modeli
 ng\,\nunderstand the impact of in-core execution on the performance of imp
 ortant algorithms from computational science such as sparse solvers\, prec
 onditioners\, and Lattice QCD\,\nunderstand the limitations of the demonst
 rated modeling approaches.\n\n\n\n\n\nCourse Structure\n\nMotivation and b
 asic processor and core architecture of the Intel Sapphire Rapids\nTermino
 logy and code execution on modern out-of-order CPUs\nx86 ISA introduction:
  understanding scalar and vectorized assembly code\nPerformance analysis o
 f simple benchmark kernels\nIntroduction to OSACA (Open-Source Architectur
 e Code Analyzer)\nIn-core analysis for ARM\, including the A64FX core arch
 itecture and AArch64 ISA\nCase Study: Sparse Matrix-Vector (SpMV) Multipli
 cation on A64FX\nCase study: Lattice Quantum Chromodynamics (QCD) on A64FX
 \nPerformance analysis\, engineering\, and optimization for a 2D Gauss-Sei
 del code on Intel Sapphire Rapids\nOutlook: Static Code Analyzers and the 
 impact of different compilers and compilation flags\n\n\n\nPrices and Elig
 ibility\nThis course is open and free of charge for participants affiliate
 d with academic institutions in European Union (EU) member states and Hori
 zon 2020-associated countries.\nRegistration\nPlease register at the botto
 m of this page. Registration is open until a few days before the course st
 arts\, or until the course is fully booked.\nWithdrawal Policy\nPlease onl
 y register if you are committed to attending the course. No-shows will be 
 blacklisted and excluded from future events.\nIf you need to withdraw your
  registration\, please either cancel it directly through the registration 
 system or send an email to nhr-training@fau.de.\nWait List\nIf the course 
 reaches its maximum capacity\, you can request to join the wait list by se
 nding an email to nhr-training@fau.de. Please include your name and univer
 sity affiliation in the message.\nAdditional Courses\nYou can find an up-t
 o-date list of all courses offered by NHR@FAU at https://hpc.fau.de/teachi
 ng/tutorials-and-courses/.\n\nhttps://indico.ecap.work/event/208/
LOCATION:Online
URL:https://indico.ecap.work/event/208/
END:VEVENT
END:VCALENDAR
