Course details

Parallel Computations on GPU

PCG Acad. year 2026/2027 Winter semester 5 credits

The course covers the architecture and programming of graphics processing units by the NVidia and partially AMD. First, the architecture of GPUs is studied in detail. Then, the model of the program execution using hierarchical thread organisation and the SIMT model is discussed. Next, the memory hierarchy and synchronization techniques are described. After that, the course explains novel techniques of dynamic parallelism and data-flow processing concluded by practical usage of multi-GPU systems in environments with shared (NVLink) and distributed (MPI) memory. The second part of the course is devoted to high level programming techniques and libraries based on the OpenACC technology.

Guarantor

Course coordinator

Language of instruction

Czech

Completion

Classified Credit (written)

Time span

  • 26 hrs lectures
  • 16 hrs pc labs
  • 20 hrs projects

Assessment points

  • 40 pts written tests (test part)
  • 60 pts projects

Department

Lecturer

Instructor

Learning objectives

To familiarize yourself with the architecture and programming of graphics processing unit in the area of general purpose computuing using the NVidia libraries and OpenACC standard. To learn how to design and implement accelerated programs exploiting the potential of GPUs. To gain knowledge about the available libraries for programming on GPUs.
Knowledge of the parallel programming on GPUs in the area of general purpose computing, orientation in the area of accelerated systems, libraries and tools.  
Understanding of hardware limitations having impact on the efficiency of software solutions. 

Prerequisite knowledge and skills

Knowledge gained in courses AVS and partially in PRL and PPP.

Study literature

  • Current PPT slides for lectures
  • Nvidia CUDA documentation: https://docs.nvidia.com/cuda/
  • OpenACC documentation: https://www.openacc.org/
  • Kirk, D., and Hwu, W.: Programming Massively Parallel Processors: A Hands-on Approach, Elsevier, 2010, s. 256, ISBN: 978-0-12-381472-2. download.
  • Sanders, J., & Kandrot, E: CUDA by Example: An Introduction to General-Purpose GPU Programming. Review Literature And Arts Of The Americas. Addison-Wesley, 2010. download.
  • Storti,D., and Yurtoglu, M.: CUDA for Engineers: An Introduction to High-Performance Parallel Computing, Addison-Wesley Professional; 1 edition, 2015. ISBN 978-0134177410. link.
  • Chandrasekaran, S., and Juckeland, G.: OpenACC for Programmers: Concepts and Strategies,  Addison-Wesley Professional, 2017, ISBN 978-0134694283. link.

Syllabus of lectures

  1. Architecture and history of graphics processing units.
  2. OpenACC library
  3. CUDA programming model, tread execution.
  4. CUDA memory hierarchy.
  5. Matrix multiplication and stencil computing
  6. Case studies of GPGPU algorithms. 
  7. Synchronization, reduction and prefix scan.
  8. Dynamic parallelism and unified memory.
  9. Stream processing, computation-communication overlapping.
  10. Multi-GPU systems.
  11. Libraries and tools for GPU programming (OpenCL, HIP, OpenMP).

Syllabus of computer exercises

  1. OpenACC: Basic Techniques (Week 2).
  2. OpenACC: Advanced Techniques (Week 3).
  3. CUDA: Memory Transfers and Simple Kernels (Week 5).
  4. CUDA: Working with Shared Memory (Week 6).
  5. CUDA: Working with Texture and Constant Memory (Week 8).
  6. CUDA: Reductions and Atomic Operations (Week 9).
  7. CUDA: Streams and CUDA Graphs (Week 11).
  8. CUDA: Dynamic Parallelism and Multi-GPU Programming (Week 12).

Syllabus - others, projects and individual work of students

  • Development of an application in OpenACC
  • Development of an application in Nvidia CUDA

Progress assessment

Assessment of two projects, 14 hours in total and, computer laboratories and a midterm examination.

  • Missed labs can be substituted in alternative dates.

Schedule

DayTypeWeeksRoomStartEndCapacityLect.grpGroupsInfo
Wed lecture 1., 2., 3., 4., 5., 6., 8., 9., 10., 11., 12., 13. of lectures E104 14:0015:5070 1MIT 2MIT xx Jaroš
Thu comp.lab lectures O204 09:0010:5016 1MIT 2MIT xx Kuník
Thu comp.lab lectures O204 11:0012:5016 1MIT 2MIT xx Kuník

Course inclusion in study plans

Back to top