StartDate: 2026-10-10 00:06:55+00:00 CpuId: 12x Intel Xeon W 2000 / D-2100 (Skylake / Cascade Lake) {Skylake}, 14nm GpuId: 1x Tesla V100-SXM2-16GB CommitSHA: d5a0442c3231244f331c9f709ad0a28469705230 CommitTime: 2026-10-09 17:07:12 +0200 CommitAuthor: Bibek Samal CommitSubject: Added fermi level computation and soc diagonalization window (#6205) #################### Building Image cp2k-perf-cuda-volta #################### Dockerfile: /tools/docker/Dockerfile.test_performance_cuda_V100 Build-Path: / Build-Args: GIT_COMMIT_SHA=d5a0442c3231244f331c9f709ad0a28469705230 SPACK_CACHE=gs://cp2k-spack-cache Build-Cache: Yes Populating docker build cache... done. DEPRECATED: The legacy builder is deprecated and will be removed in a future release. BuildKit is currently disabled; enable it by removing the DOCKER_BUILDKIT=0 environment-variable. Sending build context to Docker daemon 501.2MB Step 1/47 : FROM docker.io/nvidia/cuda:12.9.1-devel-ubuntu24.04 12.9.1-devel-ubuntu24.04: Pulling from nvidia/cuda 32f112e3802c: Pulling fs layer 644e9b203583: Pulling fs layer 02559cd4bc8d: Pulling fs layer 2cd52cbb1ebe: Pulling fs layer 6e8af4fd0a07: Pulling fs layer 15a17189b2df: Pulling fs layer 02cb0e091e33: Pulling fs layer 9c3d619183d2: Pulling fs layer 7f7602a82106: Pulling fs layer 5a2aba542b08: Pulling fs layer 6cb9b761b877: Pulling fs layer 02cb0e091e33: Waiting 9c3d619183d2: Waiting 7f7602a82106: Waiting 2cd52cbb1ebe: Waiting 5a2aba542b08: Waiting 6e8af4fd0a07: Waiting 15a17189b2df: Waiting 6cb9b761b877: Waiting 644e9b203583: Verifying Checksum 644e9b203583: Download complete 32f112e3802c: Verifying Checksum 32f112e3802c: Download complete 2cd52cbb1ebe: Verifying Checksum 2cd52cbb1ebe: Download complete 6e8af4fd0a07: Verifying Checksum 6e8af4fd0a07: Download complete 02cb0e091e33: Verifying Checksum 02cb0e091e33: Download complete 9c3d619183d2: Verifying Checksum 9c3d619183d2: Download complete 7f7602a82106: Verifying Checksum 7f7602a82106: Download complete 02559cd4bc8d: Verifying Checksum 02559cd4bc8d: Download complete 6cb9b761b877: Verifying Checksum 6cb9b761b877: Download complete 32f112e3802c: Pull complete 644e9b203583: Pull complete 02559cd4bc8d: Pull complete 2cd52cbb1ebe: Pull complete 6e8af4fd0a07: Pull complete 15a17189b2df: Verifying Checksum 15a17189b2df: Download complete 5a2aba542b08: Verifying Checksum 5a2aba542b08: Download complete 15a17189b2df: Pull complete 02cb0e091e33: Pull complete 9c3d619183d2: Pull complete 7f7602a82106: Pull complete 5a2aba542b08: Pull complete 6cb9b761b877: Pull complete Digest: sha256:020bc241a628776338f4d4053fed4c38f6f7f3d7eb5919fecb8de313bb8ba47c Status: Downloaded newer image for nvidia/cuda:12.9.1-devel-ubuntu24.04 ---> eecafe98c3e1 Step 2/47 : ENV CUDA_PATH /usr/local/cuda ---> Using cache ---> 780681fb1fee Step 3/47 : ENV LD_LIBRARY_PATH /usr/local/cuda/lib64 ---> Using cache ---> ba98a15dc225 Step 4/47 : ENV CUDA_CACHE_DISABLE 1 ---> Using cache ---> 3932740340f7 Step 5/47 : RUN apt-get update -qq && apt-get install -qq --no-install-recommends gfortran && rm -rf /var/lib/apt/lists/* ---> Using cache ---> a06eb14abc29 Step 6/47 : WORKDIR /opt/cp2k-toolchain ---> Using cache ---> 082681bac850 Step 7/47 : COPY ./tools/toolchain/install_requirements*.sh ./ ---> Using cache ---> ae920e0abda3 Step 8/47 : RUN ./install_requirements.sh ubuntu ---> Using cache ---> 94839a704e2d Step 9/47 : RUN mkdir scripts ---> Using cache ---> 433a8b0a0499 Step 10/47 : COPY ./tools/toolchain/scripts/VERSION ./tools/toolchain/scripts/tool_kit.sh ./tools/toolchain/scripts/common_vars.sh ./tools/toolchain/scripts/signal_trap.sh ./scripts/ ---> Using cache ---> edd0ada5e677 Step 11/47 : COPY ./tools/toolchain/install_cp2k_toolchain.sh . ---> fb0e7c8537de Step 12/47 : RUN ./install_cp2k_toolchain.sh --with-mpich=install --mpi-mode=mpich --enable-cuda=yes --with-libgint=install --with-sirius=install --gpu-ver=V100 --dry-run ---> Running in 5aac1aa0a420 No MPI installation detected. (Ignore this message if a fresh MPI installation is requested.) Toolchain script received the following options: --with-mpich=install --mpi-mode=mpich --enable-cuda=yes --with-libgint=install --with-sirius=install --gpu-ver=V100 --dry-run Parsing options and resolving conflicts... WARNING: (./install_cp2k_toolchain.sh, line 1128) Installing dependencies and CP2K requires CMake but CMake is not enabled, so a new copy of CMake will be installed first.  Toolchain configuration summary ------------------------------- System specifications: -j = 12 --target-cpu = native --gpu-ver = V100 --mpi-mode = mpich --math-mode = openblas Enabled features: --enable-tsan = no --enable-cuda = yes --enable-gauxc-cutlass = no --enable-hip = no --enable-opencl = no --enable-cray = no Packages to be installed: - cmake - mpich - openblas - fftw - eigen - libint - libxc - libxsmm - libxs - cosma - scalapack - elpa - dbcsr - spfft - spla - gsl - spglib - hdf5 - libvdwxc - sirius - libvori - tblite - pugixml - fmt - libgint - libwignernj Packages to be detected from system: - gcc Packages not used: - intel - amd - ninja - openmpi - intelmpi - mkl - acml - gauxc - libxstream - cusolvermp - plumed - libtorch - deepmd - ace - dftd4 - libsmeagol - trexio - libfci - greenx - gmp - mcl - skala_ftorch With --dry-run option, this script concludes with above report. The setup, toolchain env and conf files are written to /opt/cp2k-toolchain/install. ---> Removed intermediate container 5aac1aa0a420 ---> f73423fbe868 Step 13/47 : COPY ./tools/toolchain/scripts/stage0/ ./scripts/stage0/ ---> 1b1a7e72c5d3 Step 14/47 : RUN ./scripts/stage0/install_stage0.sh && rm -rf ./build ---> Running in 9303ac5c0262 ==================== Finding GCC from system paths ==================== path to gcc is /usr/bin/gcc path to g++ is /usr/bin/g++ path to gfortran is /usr/bin/gfortran GCC compiler version 13.3.0 found Step gcc took 0.00 seconds. Step intel took 0.00 seconds. Step amd took 0.00 seconds. ==================== Installing CMake ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/cmake-4.3.0-linux-x86_64.tar.gz -O cmake-4.3.0-linux-x86_64.tar.gz cmake-4.3.0-linux-x86_64.tar.gz: OK Checksum of cmake-4.3.0-linux-x86_64.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/cmake-4.3.0 Step cmake took 6.00 seconds. Step ninja took 0.00 seconds. ---> Removed intermediate container 9303ac5c0262 ---> 920b92a4492f Step 15/47 : COPY ./tools/toolchain/scripts/stage1/ ./scripts/stage1/ ---> db9eb21e8626 Step 16/47 : RUN ./scripts/stage1/install_stage1.sh && rm -rf ./build ---> Running in 8d1b2672ed1a ==================== Installing MPICH ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/mpich-5.0.2.tar.gz -O mpich-5.0.2.tar.gz mpich-5.0.2.tar.gz: OK Checksum of mpich-5.0.2.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/mpich-5.0.2 for MPICH device ch4 Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/bin Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/lib Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/include mpiexec is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpiexec mpicc is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpicc mpicxx is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpicxx mpifort is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpifort Step mpich took 647.00 seconds. ---> Removed intermediate container 8d1b2672ed1a ---> d0c31d74add1 Step 17/47 : COPY ./tools/toolchain/scripts/stage2/ ./scripts/stage2/ ---> fe9c0e782774 Step 18/47 : RUN ./scripts/stage2/install_stage2.sh && rm -rf ./build ---> Running in fc749dea023a ==================== Installing OpenBLAS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/OpenBLAS-0.3.34.tar.gz -O OpenBLAS-0.3.34.tar.gz OpenBLAS-0.3.34.tar.gz: OK Checksum of OpenBLAS-0.3.34.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/openblas-0.3.34 Installing OpenBLAS library for native target Step openblas took 331.00 seconds. Step gmp took 0.00 seconds. ---> Removed intermediate container fc749dea023a ---> f39453350d3d Step 19/47 : COPY ./tools/toolchain/scripts/stage3/ ./scripts/stage3/ ---> 92d1420701d6 Step 20/47 : RUN ./scripts/stage3/install_stage3.sh && rm -rf ./build ---> Running in 51f51e507af0 ==================== Installing FFTW ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/fftw-3.3.11.tar.gz -O fftw-3.3.11.tar.gz fftw-3.3.11.tar.gz: OK Checksum of fftw-3.3.11.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/fftw-3.3.11 Step fftw took 186.00 seconds. ==================== Installing Eigen ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/eigen-5.0.1.tar.gz -O eigen-5.0.1.tar.gz eigen-5.0.1.tar.gz: OK Checksum of eigen-5.0.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/eigen-5.0.1 Step eigen took 4.00 seconds. ==================== Installing LIBINT ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libint-v2.13.1-cp2k-lmax-5.tar.xz -O libint-v2.13.1-cp2k-lmax-5.tar.xz libint-v2.13.1-cp2k-lmax-5.tar.xz: OK Checksum of libint-v2.13.1-cp2k-lmax-5.tar.xz Ok Installing from scratch into /opt/cp2k-toolchain/install/libint-v2.13.1-cp2k-lmax-5 Step libint took 579.00 seconds. ==================== Installing LIBXC ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxc-7.1.2.tar.bz2 -O libxc-7.1.2.tar.bz2 libxc-7.1.2.tar.bz2: OK Checksum of libxc-7.1.2.tar.bz2 Ok Installing from scratch into /opt/cp2k-toolchain/install/libxc-7.1.2 Installing CUDA-only libxc into /opt/cp2k-toolchain/install/libxc-7.1.2 Step libxc took 467.00 seconds. Step greenx took 0.00 seconds. ---> Removed intermediate container 51f51e507af0 ---> d7e5fe17680c Step 21/47 : COPY ./tools/toolchain/scripts/stage4/ ./scripts/stage4/ ---> 9d85a238f672 Step 22/47 : RUN ./scripts/stage4/install_stage4.sh && rm -rf ./build ---> Running in 89039955a79d ==================== Installing Libxsmm ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxsmm-2.1.0.tar.gz -O libxsmm-2.1.0.tar.gz libxsmm-2.1.0.tar.gz: OK Checksum of libxsmm-2.1.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libxsmm-2.1.0 Step libxsmm took 26.00 seconds. ==================== Installing LIBXS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxs-1.0.0.tar.gz -O libxs-1.0.0.tar.gz libxs-1.0.0.tar.gz: OK Checksum of libxs-1.0.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libxs-1.0.0 Step libxs took 9.00 seconds. Step libxstream took 0.00 seconds. ==================== Installing libGint ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libGint-v1.tar.gz -O libGint-v1.tar.gz libGint-v1.tar.gz: OK Checksum of libGint-v1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libGint-v1 Step libGint took 137.00 seconds. ==================== Installing ScaLAPACK ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/scalapack-2.2.3.tar.gz -O scalapack-2.2.3.tar.gz scalapack-2.2.3.tar.gz: OK Checksum of scalapack-2.2.3.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/scalapack-2.2.3 Step scalapack took 42.00 seconds. Step cusolvermp took 0.00 seconds. ==================== Installing COSMA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/COSMA-v2.8.4.tar.gz -O COSMA-v2.8.4.tar.gz COSMA-v2.8.4.tar.gz: OK Checksum of COSMA-v2.8.4.tar.gz Ok wget --tries=5 --quiet https://www.cp2k.org/static/downloads/COSTA-v2.3.2.tar.gz -O COSTA-v2.3.2.tar.gz COSTA-v2.3.2.tar.gz: OK Checksum of COSTA-v2.3.2.tar.gz Ok wget --tries=5 --quiet https://www.cp2k.org/static/downloads/Tiled-MM-v2.3.2.tar.gz -O Tiled-MM-v2.3.2.tar.gz Tiled-MM-v2.3.2.tar.gz: OK Checksum of Tiled-MM-v2.3.2.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/COSMA-2.8.4 Step cosma took 76.00 seconds. ---> Removed intermediate container 89039955a79d ---> c50011bd910f Step 23/47 : COPY ./tools/toolchain/scripts/stage5/ ./scripts/stage5/ ---> dc9fa747020e Step 24/47 : RUN ./scripts/stage5/install_stage5.sh && rm -rf ./build ---> Running in 5026ddd6280b ==================== Installing ELPA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/elpa-2026.02.002.tar.gz -O elpa-2026.02.002.tar.gz elpa-2026.02.002.tar.gz: OK Checksum of elpa-2026.02.002.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002 Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002/cpu Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia Step elpa took 373.00 seconds. Step skala took 0.00 seconds. ---> Removed intermediate container 5026ddd6280b ---> e2fa23621d0f Step 25/47 : COPY ./tools/toolchain/scripts/stage6/ ./scripts/stage6/ ---> 1c9153d536dc Step 26/47 : RUN ./scripts/stage6/install_stage6.sh && rm -rf ./build ---> Running in 8c054f8fe548 ==================== Installing GSL ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/gsl-2.8.tar.gz -O gsl-2.8.tar.gz gsl-2.8.tar.gz: OK Checksum of gsl-2.8.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/gsl-2.8 Step gsl took 83.00 seconds. Step plumed took 0.00 seconds. Step libtorch took 0.00 seconds. Step ftorch took 0.00 seconds. Step gauxc took 0.00 seconds. Step deepmd took 0.00 seconds. Step ace took 0.00 seconds. ---> Removed intermediate container 8c054f8fe548 ---> 7723827207a3 Step 27/47 : COPY ./tools/toolchain/scripts/stage7/ ./scripts/stage7/ ---> 9a52b15aabdf Step 28/47 : RUN ./scripts/stage7/install_stage7.sh && rm -rf ./build ---> Running in e9fadd6bf0b8 ==================== Installing HDF5 ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/hdf5-2.2.0.tar.gz -O hdf5-2.2.0.tar.gz hdf5-2.2.0.tar.gz: OK Checksum of hdf5-2.2.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/hdf5-2.2.0 Step hdf5 took 143.00 seconds. ==================== Installing libvdwxc ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libvdwxc-0.5.0.tar.gz -O libvdwxc-0.5.0.tar.gz libvdwxc-0.5.0.tar.gz: OK Checksum of libvdwxc-0.5.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libvdwxc-0.5.0 Step libvdwxc took 17.00 seconds. ==================== Installing Spglib ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/spglib-2.7.0.tar.gz -O spglib-2.7.0.tar.gz spglib-2.7.0.tar.gz: OK Checksum of spglib-2.7.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/spglib-2.7.0 Step spglib took 5.00 seconds. ==================== Installing libvori ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libvori-220621.tar.gz -O libvori-220621.tar.gz libvori-220621.tar.gz: OK Checksum of libvori-220621.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libvori-220621 Step libvori took 15.00 seconds. ==================== Installing libwignernj ==================== wget --tries=5 --quiet https://github.com/susilehtola/libwignernj/archive/refs/tags/v0.8.0.tar.gz -O libwignernj-0.8.0.tar.gz libwignernj-0.8.0.tar.gz: OK Checksum of v0.8.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libwignernj-0.8.0 Step libwignernj took 2.00 seconds. Step libsmeagol took 0.00 seconds. Step libfci took 0.00 seconds. ==================== Installing fmt ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/fmt-12.1.0.zip -O fmt-12.1.0.zip fmt-12.1.0.zip: OK Checksum of fmt-12.1.0.zip Ok Installing from scratch into /opt/cp2k-toolchain/install/fmt-12.1.0 Step fmt took 9.00 seconds. ---> Removed intermediate container e9fadd6bf0b8 ---> 29778a8c4afa Step 29/47 : COPY ./tools/toolchain/scripts/stage8/ ./scripts/stage8/ ---> 63bc9abaf271 Step 30/47 : RUN ./scripts/stage8/install_stage8.sh && rm -rf ./build ---> Running in 66351d84b0b2 Step dftd4 took 0.00 seconds. ==================== Installing tblite ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/tblite-0.7.0.tar.xz -O tblite-0.7.0.tar.xz tblite-0.7.0.tar.xz: OK Checksum of tblite-0.7.0.tar.xz Ok Step tblite took 49.00 seconds. ==================== Installing pugixml ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/pugixml-1.15.tar.gz -O pugixml-1.15.tar.gz pugixml-1.15.tar.gz: OK Checksum of pugixml-1.15.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/pugixml-1.15 Step pugixml took 10.00 seconds. ==================== Installing SpFFT ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SpFFT-1.1.1.tar.gz -O SpFFT-1.1.1.tar.gz SpFFT-1.1.1.tar.gz: OK Checksum of SpFFT-1.1.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/SpFFT-1.1.1 Step spfft took 24.00 seconds. ==================== Installing SpLA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SpLA-1.6.1.tar.gz -O SpLA-1.6.1.tar.gz SpLA-1.6.1.tar.gz: OK Checksum of SpLA-1.6.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/SpLA-1.6.1 Step spla took 27.00 seconds. ==================== Installing SIRIUS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SIRIUS-7.11.1.tar.gz -O SIRIUS-7.11.1.tar.gz SIRIUS-7.11.1.tar.gz: OK Checksum of SIRIUS-7.11.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/sirius-7.11.1 Installing from scratch into /opt/cp2k-toolchain/install/sirius-7.11.1/cuda Step sirius took 505.00 seconds. Step trexio took 0.00 seconds. Step MCL took 0.00 seconds. ---> Removed intermediate container 66351d84b0b2 ---> 93c45e5b9b07 Step 31/47 : COPY ./tools/toolchain/scripts/stage9/ ./scripts/stage9/ ---> 39516bf138e8 Step 32/47 : RUN ./scripts/stage9/install_stage9.sh && rm -rf ./build ---> Running in 9f4abec97ca6 ==================== Installing DBCSR ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/dbcsr-2.10.0.tar.gz -O dbcsr-2.10.0.tar.gz dbcsr-2.10.0.tar.gz: OK Checksum of dbcsr-2.10.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/dbcsr-2.10.0 Installing from scratch into /opt/cp2k-toolchain/install/dbcsr-2.10.0-cuda Step DBCSR took 143.00 seconds. ---> Removed intermediate container 9f4abec97ca6 ---> 2e4557587030 Step 33/47 : WORKDIR /opt/cp2k ---> Running in 8811747c2cc0 ---> Removed intermediate container 8811747c2cc0 ---> 43bc07ea44d2 Step 34/47 : COPY ./src ./src ---> f1afd416aba3 Step 35/47 : COPY ./data ./data ---> eb1e231fe197 Step 36/47 : COPY ./tools/build_utils ./tools/build_utils ---> 6ccb1d2759ed Step 37/47 : COPY ./cmake ./cmake ---> 0cab0031f645 Step 38/47 : COPY ./CMakeLists.txt . ---> 45d2ba296b08 Step 39/47 : COPY ./CMakePresets.json . ---> 4acaf2673d8f Step 40/47 : COPY ./tools/docker/scripts/build_cp2k.sh ./tools/docker/scripts/cmake_cp2k.sh ./ ---> 9f3593a8ea27 Step 41/47 : RUN ./build_cp2k.sh toolchain_cuda_V100 psmp ---> Running in 8fdd42230c81 ==================== Building CP2K ==================== -- The Fortran compiler identification is GNU 13.3.0 -- The C compiler identification is GNU 13.3.0 -- The CXX compiler identification is GNU 13.3.0 -- Detecting Fortran compiler ABI info -- Detecting Fortran compiler ABI info - done -- Check for working Fortran compiler: /usr/bin/gfortran - skipped -- Detecting C compiler ABI info -- Detecting C compiler ABI info - done -- Check for working C compiler: /usr/bin/gcc - skipped -- Detecting C compile features -- Detecting C compile features - done -- Detecting CXX compiler ABI info -- Detecting CXX compiler ABI info - done -- Check for working CXX compiler: /usr/bin/g++ - skipped -- Detecting CXX compile features -- Detecting CXX compile features - done -- Found PkgConfig: /usr/bin/pkg-config (found version "1.8.1") -- Found Python: /usr/bin/python3.12 (found version "3.12.3") found components: Interpreter -- Found MPI_C: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpi.so (found version "5.0") -- Found MPI_CXX: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpicxx.so (found version "5.0") -- Found MPI_Fortran: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpifort.so (found version "5.0") -- Found MPI: TRUE (found version "5.0") found components: C CXX Fortran -- Could NOT find MKL (missing: CP2K_MKL_INCLUDE_DIRS _mkl_interface_library _mkl_thread_library _mkl_core_library _mkl_scalapack_library _mkl_blacs_library) -- Checking for module 'openblas' -- Found openblas, version 0.3.34 -- Found OpenBLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/include -- Found Blas: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found Lapack: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Checking for module 'scalapack' -- Package 'mpi', required by 'scalapack', not found Package 'lapack', required by 'scalapack', not found Package 'blas', required by 'scalapack', not found -- Found SCALAPACK: /opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a -- Found Threads: TRUE -- Using LIBXS + LIBXSMM for Small Matrix Multiplication -- CP2K_WITH_GPU is deprecated in favor of CMAKE_HIP_ARCHITECTURES or CMAKE_CUDA_ARCHITECTURES ------------------------------------------------------------ - DBCSR - ------------------------------------------------------------ -- Found MPI: TRUE (found version "5.0") -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") -- The CUDA compiler identification is NVIDIA 12.9.86 with host compiler GNU 13.3.0 -- Detecting CUDA compiler ABI info -- Detecting CUDA compiler ABI info - done -- Check for working CUDA compiler: /usr/local/cuda/bin/nvcc - skipped -- Detecting CUDA compile features -- Detecting CUDA compile features - done -- Found CUDAToolkit: /usr/local/cuda/targets/x86_64-linux/include (found version "12.9.86") ----------------------------------------------------------- - CUDA - ----------------------------------------------------------- -- GPU architecture number: 70 -- GPU profiling enabled: OFF -- CUDA compiler and libraries found ------------------------------------------------------------ - OPENMP - ------------------------------------------------------------ -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") found components: Fortran C CXX ------------------------------------------------------------ - Other dependencies - ------------------------------------------------------------ -- Checking for one of the modules 'elpa_openmp' -- Found Elpa: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so;cudart;cublasLt;cublas;/opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a;:libopenblas.a -- Found CUDA Libxc (7.1.2) -- Found HDF5: hdf5-shared;hdf5_fortran-shared (found version "2.2.0") found components: C Fortran -- Found MPI: TRUE (found version "5.0") found components: CXX -- Found OPENBLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found Blas: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Checking for one of the modules 'fftw3' -- Checking for one of the modules 'fftw3f' -- Checking for one of the modules 'fftw3l' -- Checking for one of the modules 'fftw3q' -- Found Fftw: /opt/cp2k-toolchain/install/fftw-3.3.11/include -- Boost detected. satisfied by headers bundled with Libint2 distribution -- Found LibGint: /opt/cp2k-toolchain/install/libGint-v1/lib/libcp2kGint.a -- Component omp of Spglib: NOT FOUND -- Component fortran of Spglib: FOUND (LIB_TYPE: static) -- Found package: Spglib -- Looking for Fortran sgemm -- Looking for Fortran sgemm - found -- multicharge: Find installed package -- toml-f: Find installed package -- s-dftd3: Find installed package -- Found GSL: /opt/cp2k-toolchain/install/gsl-2.8/include (found version "2.8") -- Checking for one of the modules 'libxc>=3.0.0' -- Found LibXC: /opt/cp2k-toolchain/install/libxc-7.1.2/lib/libxc.so (Required is at least version "3.0.0") -- Found LibSPG: /opt/cp2k-toolchain/install/spglib-2.7.0/lib/libsymspg.a -- Found HDF5: hdf5-shared (found version "2.2.0") found components: C -- Found FFTW: /opt/cp2k-toolchain/install/fftw-3.3.11/include -- Looking for Fortran sgemm -- Looking for Fortran sgemm - not found -- Found BLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP_CUDA: -fopenmp (found version "4.5") -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") -- Checking for one of the modules 's-dftd3' -- Checking for one of the modules 'mctc-lib' -- Found DFTD3: /opt/cp2k-toolchain/install/tblite-0.7.0/lib/libs-dftd3.a -- Checking for one of the modules 'dftd4' -- Checking for one of the modules 'multicharge' -- Found DFTD4: /opt/cp2k-toolchain/install/tblite-0.7.0/lib/libdftd4.a -- Looking for Fortran cheev -- Looking for Fortran cheev - found -- Found LAPACK: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so;-lm;-ldl -- Checking for one of the modules 'scalapack' -- Checking for one of the modules 'elpa;elpa_openmp;elpa-openmp-2019.05.001;elpa_openmp-2019.11.001;elpa_openmp-2020.05.001;elpa-2019.05.001;elpa-2019.11.001;elpa-2020.05.001' -- Found Elpa: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so -- Checking for module 'libvdwxc>=0.5.0' -- Found libvdwxc, version 0.5.0 -- Checking for module 'fftw3' -- Found fftw3, version 3.3.11 -- Found LibVDWXC: vdwxc;fftw3 (Required is at least version "0.5.0") -- Setting build type to 'Release' as none was specified. -- Performing Test f2008-norm2 -- Performing Test f2008-norm2 - Success -- Performing Test f2008-block_construct -- Performing Test f2008-block_construct - Success -- Performing Test f2008-contiguous -- Performing Test f2008-contiguous - Success -- Performing Test f2008-findloc -- Performing Test f2008-findloc - Success -- Performing Test f95-reshape-order-allocatable -- Performing Test f95-reshape-order-allocatable - Success -- FYPP preprocessor found. -- Adding libxs_jit.F from dependency libxs for compilation -------------------------------------------------------------------- - - - Summary of enabled dependencies - - - -------------------------------------------------------------------- - BLAS - Vendor: OpenBLAS - Include directories: /opt/cp2k-toolchain/install/openblas-0.3.34/include - Libraries: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so - LAPACK - Include directories: /opt/cp2k-toolchain/install/openblas-0.3.34/include - Libraries: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so - FFTW3 - Include directories: /opt/cp2k-toolchain/install/fftw-3.3.11/include - Libraries: /opt/cp2k-toolchain/install/fftw-3.3.11/lib/libfftw3.a - MPI - Include directories: /opt/cp2k-toolchain/install/mpich-5.0.2/include - Libraries: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpicxx.so;/opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpi.so - MPI_F08: Enabled - ScaLAPACK - Vendor: auto - Include directories: - Libraries: /opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a - Hardware acceleration - Backend: CUDA - GPU architectures: 70 - GPU profiling enabled: OFF - GPU-accelerated modules - ELPA: ON - GRID: ON - DBM: ON - PW: ON - LIBXC: ON - LibXC - Include directories: LIBXC_CPU_INCLUDE_DIRS-NOTFOUND - Libraries: Libxc::xc - Spglib - Include directories: /opt/cp2k-toolchain/install/spglib-2.7.0/include;$ - HDF5 - Include directories: /opt/cp2k-toolchain/install/hdf5-2.2.0/include - Libraries: hdf5-shared - LIBXS - Include directories: - Libraries: - SpLA - Include directories: /opt/cp2k-toolchain/install/SpLA-1.6.1-cuda/include;/opt/cp2k-toolchain/install/SpLA-1.6.1-cuda/include/spla - Libraries: $;$;$;$;MPI::MPI_CXX;MPI::MPI_C;MPI::MPI_Fortran - SpLA GEMM offloading - DFTD4 - Enabled via TBLITE - Include directories: /opt/cp2k-toolchain/install/tblite-0.7.0/include;/opt/cp2k-toolchain/install/tblite-0.7.0/include/dftd4/GNU-13.3.0 - Libraries: - TBLITE - Include directories: - Libraries: - SIRIUS - Include directories: - Libraries: - COSMA - Include directories: /opt/cp2k-toolchain/install/COSMA-2.8.4-cuda/include - Libraries: MPI::MPI_CXX;costa::costa;$;$;$<$:cosma::BLAS::blas>;$;$<$:Tiled-MM::Tiled-MM>;$<$:Tiled-MM::Tiled-MM>;$<$:semiprof::semiprof>;$<$:cosma::scalapack::scalapack> - Libint2 - Include directories: - Libraries: - LibGint - include directories: /opt/cp2k-toolchain/install/libGint-v1/include - libraries: /opt/cp2k-toolchain/install/libGint-v1/lib/libcp2kGint.a - Libwignernj - Version: 0.8.0 - Libraries: wignernj::wignernj - ELPA - Include directories: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/include/elpa_openmp-2026.02.002 - Libraries: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so;cudart;cublasLt;cublas;/opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a;:libopenblas.a -------------------------------------------------------------------- - - - Dependencies not included in this build - - - -------------------------------------------------------------------- - DeePMD - PEXSI - ACE (libpace) - LibSMEAGOL - MiMiC - DLA-Future - PLUMED - LibFCI - GauXC - Skala/FTorch - Libvori - LibTorch - TREXIO - OpenPMD - GreenX After building and installing CP2K, run the regtests with: /opt/cp2k/tests/do_regtest.py /opt/cp2k/bin psmp -- Configuring done (15.6s) -- Generating done (1.0s) -- Build files have been written to: /opt/cp2k/build Compiling CP2K ... done ---> Removed intermediate container 8fdd42230c81 ---> 11f33a92eba6 Step 42/47 : COPY ./benchmarks ./benchmarks ---> 4c37debac91c Step 43/47 : COPY ./tools/regtesting ./tools/regtesting ---> 8fbcd3ad0d84 Step 44/47 : COPY ./tools/docker/scripts/test_performance.sh ./tools/docker/scripts/plot_performance.py ./ ---> 8127353d69e5 Step 45/47 : RUN ./test_performance.sh "toolchain_cuda_V100" 2>&1 | tee report.log ---> Running in f1a5e1bc73cd ============== CP2K Binary Flags ============= cp2kflags: omp libint fftw3 libxc libxc_gpu elpa parallel scalapack mpi_f08 cosma libxs libxsmm dbcsr_acc spglib openblas libdftd4 s_dftd3 mctc-lib tblite sirius offload_cuda spla_gemm_offloading libvdwxc hdf5 libGint libwignernj ========== Checking Benchmark Inputs ========= Found 89 input files and 0 errors. ========== Running Performance Test ========== Plot: name="total_timings_6cpu_1gpu", title="Total Timings with 6 CPU Cores and 1 GPU", ylabel="time [s]" Running H2O-64.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/H2O-64_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.033 0.034 112.657 112.657 qs_mol_dyn_low 1 2.0 0.005 0.005 112.182 112.185 qs_forces 11 3.9 0.002 0.002 112.127 112.127 qs_energies 11 4.9 0.002 0.002 99.552 99.553 scf_env_do_scf 11 5.9 0.001 0.001 82.668 82.669 velocity_verlet 10 3.0 0.002 0.002 70.413 70.432 scf_env_do_scf_inner_loop 108 6.5 0.007 0.009 67.808 67.808 rebuild_ks_matrix 119 8.3 0.001 0.001 29.040 29.040 qs_ks_build_kohn_sham_matrix 119 9.3 0.028 0.028 29.039 29.039 dbcsr_multiply_generic 2319 12.5 0.170 0.172 27.775 27.857 qs_ks_update_qs_env 119 7.6 0.001 0.002 26.963 26.965 qs_rho_update_rho_low 119 7.7 0.001 0.001 23.217 23.236 calculate_rho_elec 119 8.7 0.949 0.961 23.216 23.235 qs_scf_new_mos 108 7.5 0.001 0.001 22.980 22.990 qs_scf_loop_do_ot 108 8.5 0.001 0.001 22.979 22.989 ot_scf_mini 108 9.5 0.003 0.003 20.846 20.848 fft_wrap_pw1pw2 1201 11.6 0.025 0.025 17.926 17.962 fft_wrap_pw1pw2_140 487 12.2 0.003 0.003 15.418 15.496 sum_up_and_integrate 119 10.3 0.005 0.005 15.064 15.160 integrate_v_rspace 119 11.3 0.378 0.383 14.942 15.038 init_scf_loop 11 6.9 0.001 0.001 14.791 14.792 multiply_cannon 2319 13.5 0.390 0.395 13.812 13.829 multiply_cannon_loop 2319 14.5 0.303 0.310 12.586 12.596 ot_mini 108 10.5 0.001 0.001 12.108 12.109 make_m2s 4638 13.5 0.051 0.052 12.069 12.094 density_rs2pw 119 9.7 0.009 0.009 11.812 11.924 make_images 4638 14.5 1.234 1.235 11.873 11.895 prepare_preconditioner 11 7.9 0.000 0.000 11.249 11.253 make_preconditioner 11 8.9 0.000 0.000 11.249 11.253 grid_collocate_task_list 119 9.7 10.419 10.497 10.419 10.497 make_full_inverse_cholesky 11 9.9 0.003 0.003 10.077 10.362 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 9.199 9.276 pw_gpu_r3dc1d_3d_ps 606 13.1 2.539 2.573 9.197 9.228 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 8.755 8.755 pw_gpu_c1dr3d_3d_ps 595 14.2 2.401 2.433 8.697 8.703 grid_integrate_task_list 119 12.3 7.616 7.716 7.616 7.716 qs_ot_get_derivative 108 11.5 0.002 0.002 7.515 7.516 init_scf_run 11 5.9 0.000 0.000 7.414 7.414 scf_env_initial_rho_setup 11 6.9 0.001 0.001 7.413 7.413 potential_pw2rs 119 12.3 0.040 0.041 6.947 6.948 make_images_data 4638 15.5 0.066 0.067 6.927 6.938 multiply_cannon_multrec 4638 15.5 2.123 2.210 6.813 6.836 hybrid_alltoall_any 4638 16.5 5.116 5.158 6.666 6.676 copy_dbcsr_to_fm 153 11.3 0.152 0.153 5.813 5.817 transfer_dbcsr_to_fm 11 10.9 0.028 0.028 5.363 5.365 dbcsr_to_fm_plan_create 11 12.9 4.506 4.640 5.022 5.025 mp_alltoall_z22v 1201 15.6 4.681 4.814 4.681 4.814 build_core_hamiltonian_matrix 11 6.9 0.002 0.002 4.588 4.604 ot_diis_step 108 11.5 0.007 0.007 4.568 4.568 build_core_ppl_forces 11 5.9 4.389 4.447 4.389 4.447 dbcsr_mm_accdrv_process 9638 16.2 1.410 2.148 4.265 4.317 wfi_extrapolate 11 7.9 0.002 0.002 4.271 4.271 mp_waitall_1 65419 16.9 4.050 4.191 4.050 4.191 apply_preconditioner_dbcsr 119 12.6 0.000 0.000 3.937 3.938 apply_single 119 13.6 0.001 0.001 3.936 3.937 qs_ot_get_p 119 10.4 0.002 0.002 3.860 3.861 calculate_dm_sparse 119 9.5 0.001 0.001 3.723 3.737 qs_env_update_s_mstruct 11 6.9 0.000 0.000 3.690 3.702 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 3.284 3.284 multiply_cannon_sync_h2d 4638 15.5 3.080 3.172 3.080 3.172 qs_ot_get_derivative_taylor 59 13.0 0.003 0.003 3.090 3.093 pw_poisson_solve 119 10.3 0.004 0.004 3.003 3.006 transfer_rs2pw 487 10.6 0.010 0.010 2.927 3.004 yz_to_x 606 14.1 0.525 0.527 2.901 2.961 x_to_yz 595 15.2 0.567 0.570 2.871 2.938 cp_dbcsr_sm_fm_multiply 37 9.5 0.002 0.002 2.919 2.919 jit_kernel_multiply 12 16.0 2.205 2.898 2.205 2.898 build_kinetic_matrix_low 22 6.9 2.685 2.701 2.795 2.810 qs_create_task_list 11 7.9 0.000 0.000 2.764 2.793 generate_qs_task_list 11 8.9 1.279 1.282 2.764 2.793 cp_fm_cholesky_invert 11 10.9 2.667 2.667 2.667 2.667 calculate_first_density_matrix 1 7.0 0.000 0.000 2.642 2.642 transfer_rs2pw_140 130 11.5 1.764 1.793 2.441 2.535 qs_ot_p2m_diag 50 11.0 0.093 0.095 2.502 2.505 cp_dbcsr_sm_fm_multiply_core 37 10.5 0.000 0.000 2.434 2.434 dbcsr_special_finalize 6957 15.5 0.046 0.047 2.352 2.364 pw_gpu_fg 606 14.1 2.319 2.364 2.319 2.364 dbcsr_complete_redistribute 318 12.2 0.795 0.808 2.014 2.295 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="H2O-64", label="H2O-64", y=112.657, yerr=0.0 Plot: name="H2O-64_timings_6cpu_1gpu", title="Timings of H2O-64 with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="rest", label="rest", y=80.31899999999999, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=10.419, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=7.616, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=5.116, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=4.681, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=4.506, yerr=0.0 Running H2O-64_nonortho.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/H2O-64_nonortho_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.031 0.033 104.145 104.145 qs_mol_dyn_low 1 2.0 0.006 0.006 103.666 103.668 qs_forces 11 3.9 0.002 0.003 103.612 103.613 qs_energies 11 4.9 0.002 0.002 90.921 90.922 scf_env_do_scf 11 5.9 0.001 0.001 73.500 73.500 velocity_verlet 10 3.0 0.002 0.002 67.170 67.188 scf_env_do_scf_inner_loop 96 6.5 0.006 0.009 58.437 58.437 rebuild_ks_matrix 107 8.3 0.001 0.001 26.539 26.542 qs_ks_build_kohn_sham_matrix 107 9.3 0.025 0.025 26.538 26.541 dbcsr_multiply_generic 1999 12.5 0.147 0.147 25.913 25.943 qs_ks_update_qs_env 107 7.6 0.001 0.001 24.265 24.267 qs_scf_new_mos 96 7.5 0.001 0.001 20.792 20.794 qs_scf_loop_do_ot 96 8.5 0.001 0.001 20.791 20.793 ot_scf_mini 96 9.5 0.003 0.003 18.932 18.933 qs_rho_update_rho_low 107 7.7 0.001 0.001 18.460 18.471 calculate_rho_elec 107 8.7 0.842 0.849 18.459 18.470 fft_wrap_pw1pw2 1081 11.6 0.022 0.022 16.081 16.117 init_scf_loop 11 6.9 0.001 0.001 14.999 15.000 sum_up_and_integrate 107 10.3 0.005 0.005 13.913 13.931 fft_wrap_pw1pw2_140 439 12.2 0.003 0.003 13.817 13.873 integrate_v_rspace 107 11.3 0.341 0.341 13.805 13.824 multiply_cannon 1999 13.5 0.340 0.342 12.686 12.891 make_m2s 3998 13.5 0.046 0.046 11.577 11.801 make_images 3998 14.5 1.201 1.309 11.403 11.625 multiply_cannon_loop 1999 14.5 0.263 0.266 11.423 11.429 prepare_preconditioner 11 7.9 0.000 0.000 11.410 11.422 make_preconditioner 11 8.9 0.000 0.000 11.410 11.422 ot_mini 96 10.5 0.001 0.001 11.065 11.067 density_rs2pw 107 9.7 0.008 0.008 10.578 10.706 make_full_inverse_cholesky 11 9.9 0.002 0.002 10.160 10.445 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 9.620 9.621 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.002 9.164 9.266 pw_gpu_r3dc1d_3d_ps 546 13.1 2.298 2.350 8.294 8.296 pw_gpu_c1dr3d_3d_ps 535 14.2 2.150 2.178 7.760 7.793 grid_integrate_task_list 107 12.3 7.263 7.279 7.263 7.279 grid_collocate_task_list 107 9.7 7.009 7.111 7.009 7.111 init_scf_run 11 5.9 0.000 0.000 7.101 7.101 scf_env_initial_rho_setup 11 6.9 0.000 0.001 7.101 7.101 make_images_data 3998 15.5 0.055 0.056 6.763 6.802 qs_ot_get_derivative 96 11.5 0.002 0.002 6.763 6.766 hybrid_alltoall_any 3998 16.5 4.805 5.043 6.539 6.579 multiply_cannon_multrec 3998 15.5 1.864 1.874 6.316 6.329 potential_pw2rs 107 12.3 0.035 0.036 6.201 6.203 copy_dbcsr_to_fm 147 11.2 0.164 0.181 5.915 5.934 transfer_dbcsr_to_fm 11 10.9 0.040 0.051 5.499 5.526 dbcsr_to_fm_plan_create 11 12.9 4.501 4.683 5.101 5.122 qs_env_update_s_mstruct 11 6.9 0.000 0.000 4.623 4.785 build_core_hamiltonian_matrix 11 6.9 0.002 0.002 4.550 4.621 build_core_ppl_forces 11 5.9 4.338 4.435 4.338 4.435 ot_diis_step 96 11.5 0.006 0.006 4.280 4.280 mp_alltoall_z22v 1081 15.6 4.129 4.228 4.129 4.228 mp_waitall_1 56411 16.9 3.886 4.217 3.886 4.217 dbcsr_mm_accdrv_process 8494 16.1 0.828 1.005 4.080 4.102 wfi_extrapolate 11 7.9 0.001 0.001 4.079 4.079 apply_preconditioner_dbcsr 107 12.6 0.000 0.000 3.868 3.872 apply_single 107 13.6 0.001 0.001 3.867 3.872 qs_create_task_list 11 7.9 0.000 0.000 3.673 3.766 generate_qs_task_list 11 8.9 1.565 1.576 3.673 3.765 calculate_dm_sparse 107 9.5 0.001 0.001 3.373 3.377 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 3.329 3.329 qs_ot_get_p 107 10.4 0.001 0.001 3.298 3.300 cp_dbcsr_sm_fm_multiply 37 9.5 0.002 0.002 2.967 2.968 build_kinetic_matrix_low 22 6.9 2.727 2.750 2.837 2.860 transfer_rs2pw 439 10.6 0.008 0.009 2.667 2.851 jit_kernel_multiply 12 15.8 2.675 2.833 2.675 2.833 multiply_cannon_sync_h2d 3998 15.5 2.776 2.827 2.776 2.827 qs_ot_get_derivative_taylor 53 13.0 0.003 0.003 2.705 2.706 pw_poisson_solve 107 10.3 0.003 0.003 2.662 2.668 yz_to_x 546 14.1 0.467 0.472 2.584 2.643 cp_fm_cholesky_invert 11 10.9 2.568 2.568 2.568 2.568 calculate_first_density_matrix 1 7.0 0.000 0.000 2.557 2.558 x_to_yz 535 15.2 0.496 0.501 2.508 2.539 cp_dbcsr_sm_fm_multiply_core 37 10.5 0.000 0.000 2.489 2.489 transfer_rs2pw_140 118 11.5 1.577 1.583 2.245 2.439 dbcsr_complete_redistribute 306 12.1 0.762 0.802 2.057 2.336 build_core_ppl 11 7.9 2.203 2.259 2.203 2.259 build_overlap_matrix_low 22 6.9 2.062 2.082 2.172 2.192 qs_ot_p2m_diag 44 11.0 0.079 0.080 2.151 2.152 copy_fm_to_dbcsr 170 11.1 0.002 0.002 1.847 2.132 dbcsr_special_finalize 5997 15.5 0.041 0.041 2.096 2.098 pw_gpu_fg 546 14.1 2.096 2.097 2.096 2.097 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="H2O-64_nonortho", label="H2O-64_nonortho", y=104.145, yerr=0.0 Plot: name="H2O-64_nonortho_timings_6cpu_1gpu", title="Timings of H2O-64_nonortho with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="rest", label="rest", y=76.229, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=7.263, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=7.009, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=4.805, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=4.501, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="build_core_ppl_forces", label="build_core_ppl_forces", y=4.338, yerr=0.0 Running w64PBE.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/w64PBE_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.054 0.057 258.118 258.118 qs_mol_dyn_low 1 2.0 0.005 0.005 257.326 257.329 qs_forces 11 3.9 0.003 0.003 257.270 257.270 qs_energies 11 4.9 0.002 0.002 222.404 222.404 velocity_verlet 10 3.0 0.002 0.002 204.133 204.153 scf_env_do_scf 11 5.9 0.001 0.002 199.623 199.624 scf_env_do_scf_inner_loop 106 6.8 0.007 0.009 172.043 172.043 rebuild_ks_matrix 117 8.5 0.001 0.001 129.074 129.074 qs_ks_build_kohn_sham_matrix 117 9.5 0.030 0.030 129.073 129.073 qs_ks_update_qs_env 120 7.8 0.001 0.002 114.667 114.669 fft_wrap_pw1pw2 2000 12.9 0.054 0.054 75.897 75.903 fft_wrap_pw1pw2_200 1298 14.3 0.010 0.010 71.923 71.937 qs_vxc_create 117 10.5 0.003 0.003 66.946 66.983 xc_vxc_pw_create 117 11.5 1.641 1.655 66.943 66.980 qs_rho_update_rho_low 117 7.9 0.001 0.001 63.905 63.916 calculate_rho_elec 117 8.9 1.324 1.326 63.905 63.915 sum_up_and_integrate 117 10.5 0.006 0.006 46.059 46.102 integrate_v_rspace 117 11.5 0.236 0.237 45.808 45.851 xc_pw_derive 702 13.5 0.012 0.013 42.792 42.878 grid_collocate_task_list 117 9.9 42.324 42.432 42.324 42.432 pw_gpu_c1dr3d_3d_ps 1053 15.2 11.573 11.593 40.641 40.697 xc_rho_set_and_dset_create 117 12.5 1.049 1.053 36.377 36.414 pw_gpu_r3dc1d_3d_ps 947 14.5 10.448 10.535 35.188 35.250 grid_integrate_task_list 117 12.5 33.589 33.645 33.589 33.645 xc_pw_divergence 117 12.5 0.007 0.008 28.442 28.494 init_scf_loop 14 6.8 0.001 0.001 27.527 27.527 density_rs2pw 117 9.9 0.011 0.011 20.221 20.353 dbcsr_multiply_generic 2077 12.5 0.166 0.167 20.066 20.122 mp_alltoall_z22v 2000 16.9 19.983 19.986 19.983 19.986 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 19.345 19.627 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 15.226 15.227 qs_scf_new_mos 106 7.8 0.001 0.001 15.027 15.028 qs_scf_loop_do_ot 106 8.8 0.001 0.001 15.026 15.028 x_to_yz 1053 16.2 3.124 3.125 13.776 13.795 ot_scf_mini 106 9.8 0.003 0.003 13.498 13.500 xc_functional_eval 117 13.5 0.002 0.002 12.992 13.044 pbe_lda_eval 117 14.5 12.990 13.042 12.990 13.042 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 12.839 12.839 potential_pw2rs 117 12.5 0.065 0.066 11.983 11.997 yz_to_x 947 15.5 2.211 2.219 11.542 11.573 prepare_preconditioner 14 7.8 0.000 0.000 11.555 11.558 make_preconditioner 14 8.8 0.000 0.000 11.555 11.558 multiply_cannon 2077 13.5 0.367 0.371 9.873 9.885 build_core_ppl_forces 11 5.9 9.489 9.780 9.489 9.780 init_scf_run 11 5.9 0.000 0.000 9.409 9.409 scf_env_initial_rho_setup 11 6.9 0.000 0.001 9.408 9.408 pw_gpu_sf 1053 16.2 8.735 8.737 8.735 8.737 multiply_cannon_loop 2077 14.5 0.279 0.279 8.717 8.718 make_m2s 4154 13.5 0.050 0.050 8.501 8.510 build_core_hamiltonian_matrix 11 6.9 0.002 0.002 8.255 8.364 pw_gpu_fg 947 15.5 8.337 8.342 8.337 8.342 make_images 4154 14.5 1.094 1.106 8.301 8.311 ot_mini 106 10.8 0.001 0.001 8.261 8.262 wfi_extrapolate 11 7.9 0.002 0.002 7.334 7.334 build_kinetic_matrix_low 22 6.9 6.674 6.684 6.773 6.783 pw_gpu_ffc 1053 16.2 6.535 6.588 6.535 6.588 make_full_inverse_cholesky 14 9.8 0.000 0.001 6.210 6.376 build_overlap_matrix_low 22 6.9 5.723 5.736 5.820 5.833 pw_poisson_solve 117 10.5 0.004 0.004 5.506 5.513 qs_ot_get_derivative 106 11.8 0.002 0.002 5.204 5.206 transfer_rs2pw 479 10.8 0.010 0.011 4.913 5.140 pw_derive 1053 13.8 4.953 4.975 4.953 4.975 pw_gpu_cff 947 15.5 4.793 4.794 4.793 4.794 make_full_single_inverse 14 9.8 0.002 0.002 4.690 4.691 multiply_cannon_multrec 4154 15.5 1.838 1.852 4.546 4.563 make_images_data 4154 15.5 0.064 0.064 4.376 4.394 transfer_rs2pw_200 128 11.7 2.957 2.969 4.106 4.338 qs_env_update_s_mstruct 11 6.9 0.000 0.000 4.152 4.249 hybrid_alltoall_any 4154 16.5 3.045 3.049 4.116 4.133 copy_dbcsr_to_fm 143 10.8 0.098 0.099 3.869 3.876 build_core_ppl 11 7.9 3.672 3.772 3.672 3.772 mp_waitall_1 58635 17.0 3.733 3.757 3.733 3.757 transfer_dbcsr_to_fm 14 10.8 0.004 0.006 3.472 3.475 transfer_pw2rs 479 13.4 0.008 0.008 3.414 3.416 pw_copy 1755 13.0 3.304 3.308 3.304 3.308 dbcsr_to_fm_plan_create 14 12.8 2.843 2.864 3.266 3.267 ot_diis_step 106 11.8 0.006 0.006 3.036 3.036 arnoldi_generalized_ev 14 10.8 0.000 0.000 2.957 2.958 dbcsr_sym_matrix_vector_mult 1269 12.5 0.041 0.042 2.899 2.900 fft_wrap_pw1pw2_70 234 13.2 0.002 0.002 2.895 2.898 transfer_pw2rs_200 128 14.1 1.761 1.776 2.735 2.738 gev_build_subspace 23 11.5 0.013 0.013 2.722 2.722 qs_create_task_list 11 7.9 0.000 0.000 2.681 2.691 generate_qs_task_list 11 8.9 1.480 1.492 2.680 2.691 pw_poisson_set 118 11.5 0.007 0.007 2.622 2.630 apply_preconditioner_dbcsr 120 12.8 0.000 0.000 2.624 2.625 apply_single 120 13.8 0.001 0.001 2.624 2.624 dbcsr_sym_matrix_vector_mult_l 1269 13.5 2.486 2.504 2.493 2.511 qs_ot_get_derivative_taylor 89 12.9 0.005 0.005 2.419 2.424 dbcsr_mm_accdrv_process 9444 16.2 0.969 1.100 2.396 2.399 calculate_dm_sparse 117 9.7 0.001 0.001 2.280 2.285 cp_dbcsr_sm_fm_multiply 46 9.3 0.002 0.002 2.034 2.035 pw_integral_ab_c1d_c1d_gs 117 11.5 1.944 1.944 1.967 1.968 pw_axpy 1170 12.0 1.947 1.948 1.947 1.948 multiply_cannon_sync_h2d 4154 15.5 1.889 1.926 1.889 1.926 qs_ot_get_p 120 10.5 0.001 0.002 1.877 1.878 dbcsr_special_finalize 6231 15.5 0.040 0.041 1.787 1.791 dbcsr_merge_single_wm 4154 16.5 0.153 0.155 1.658 1.663 dbcsr_complete_redistribute 309 11.8 0.601 0.620 1.490 1.655 mp_sendrecv_dv 479 12.8 1.391 1.604 1.391 1.604 cp_dbcsr_sm_fm_multiply_core 46 10.3 0.000 0.000 1.575 1.575 cp_fm_cholesky_invert 14 10.8 1.545 1.545 1.545 1.545 calculate_rho_core 11 7.9 0.178 0.178 1.414 1.522 copy_fm_to_dbcsr 180 10.8 0.002 0.002 1.281 1.447 multiply_cannon_metrocomm1 4154 15.5 0.014 0.014 1.368 1.383 dbcsr_dot 1125 12.2 1.234 1.236 1.307 1.313 calculate_first_density_matrix 1 7.0 0.000 0.000 1.222 1.222 dbcsr_sort_data 4154 17.5 1.199 1.201 1.199 1.201 jit_kernel_multiply 10 15.0 0.878 1.015 0.878 1.015 cp_dbcsr_plus_fm_fm_t 22 8.9 0.001 0.001 1.013 1.014 qs_ot_get_orbitals 106 10.8 0.001 0.001 0.914 0.915 dbcsr_copy 7924 13.4 0.241 0.242 0.897 0.901 qs_ot_p2m_diag 19 11.0 0.037 0.038 0.894 0.895 build_core_ppnl_forces 11 5.9 0.884 0.888 0.884 0.888 grid_create_task_list 11 9.9 0.842 0.853 0.842 0.853 evaluate_core_matrix_traces 117 8.5 0.001 0.001 0.818 0.819 calculate_ptrace_kp 234 9.5 0.001 0.001 0.817 0.818 transfer_fm_to_dbcsr 14 9.8 0.000 0.000 0.654 0.817 mp_sum_d 3821 11.6 0.546 0.811 0.546 0.811 cp_dbcsr_syevd 19 12.0 0.002 0.002 0.763 0.763 make_images_pack 4154 15.5 0.740 0.744 0.756 0.760 fft_wrap_pw1pw2_30 234 13.2 0.001 0.001 0.751 0.758 cp_fm_cholesky_decompose 28 10.5 0.737 0.741 0.737 0.741 cp_fm_diag_elpa 19 13.0 0.000 0.000 0.728 0.729 cp_fm_diag_elpa_base 19 14.0 0.717 0.719 0.727 0.728 dbcsr_finalize 4569 14.0 0.066 0.066 0.711 0.714 qs_init_subsys 1 2.0 0.001 0.001 0.691 0.691 qs_env_setup 1 3.0 0.000 0.000 0.683 0.684 qs_env_rebuild_pw_env 23 5.3 0.000 0.000 0.683 0.683 pw_env_rebuild 1 5.0 0.000 0.000 0.682 0.683 cp_fm_uplo_to_full 47 13.4 0.489 0.659 0.489 0.659 pw_grid_setup 4 6.0 0.000 0.000 0.655 0.656 pw_zero 585 13.0 0.654 0.656 0.654 0.656 pw_grid_setup_internal 4 7.0 0.007 0.008 0.644 0.644 transfer_rs2pw_70 117 11.9 0.441 0.443 0.629 0.633 calculate_ecore_overlap 22 5.9 0.002 0.002 0.328 0.619 dbcsr_merge_all 4154 15.2 0.197 0.197 0.606 0.609 dbcsr_copy_into_existing 22 7.9 0.589 0.593 0.589 0.593 qs_ot_get_derivative_diag 17 12.0 0.001 0.001 0.585 0.586 make_basis_sm 14 9.3 0.001 0.001 0.578 0.579 acc_transpose_blocks 4154 15.5 0.026 0.026 0.573 0.577 dbcsr_mm_accdrv_process_sort 9444 17.2 0.549 0.554 0.549 0.554 mp_alltoall_d11v 1526 13.9 0.524 0.537 0.524 0.537 transfer_pw2rs_70 117 14.5 0.345 0.346 0.526 0.527 pw_grid_sort 4 8.0 0.383 0.384 0.520 0.521 dbcsr_sort_indices 10944 16.6 0.463 0.464 0.463 0.464 parallel_gemm_fm_cosma 96 8.9 0.462 0.463 0.462 0.463 ot_scf_init 14 7.8 0.002 0.002 0.440 0.443 compute_matrix_w 11 5.9 0.000 0.000 0.435 0.436 calculate_w_matrix_ot 11 6.9 0.003 0.003 0.435 0.436 reorthogonalize_vectors 10 9.0 0.000 0.000 0.414 0.414 mp_sum_l 6260 13.5 0.363 0.414 0.363 0.414 build_qs_neighbor_lists 11 6.9 0.001 0.001 0.388 0.390 cp_dbcsr_alloc_block_from_nbl 88 7.7 0.257 0.257 0.387 0.389 pw_scale 468 12.0 0.379 0.383 0.379 0.383 mp_alltoall_i22 504 14.0 0.216 0.378 0.216 0.378 integrate_v_core_rspace 11 7.9 0.078 0.078 0.363 0.366 dbcsr_add_d 1879 13.1 0.004 0.004 0.354 0.355 dbcsr_add_anytype 1879 14.1 0.190 0.193 0.351 0.352 distribute_tasks 11 9.9 0.341 0.343 0.341 0.343 mp_alltoall_i 14 13.8 0.312 0.339 0.312 0.339 multiply_cannon_multrec_finali 2077 16.5 0.006 0.006 0.312 0.313 setup_rec_index_2d 4154 14.5 0.310 0.313 0.310 0.313 dbcsr_mm_multrec_finalize 2077 17.5 0.026 0.027 0.306 0.307 pw_multiply_with 117 11.5 0.306 0.306 0.306 0.306 dbcsr_mm_sched_finalize 2077 18.5 0.274 0.276 0.280 0.282 fft_wrap_pw1pw2_10 234 13.2 0.001 0.001 0.274 0.276 dbcsr_make_untransposed_blocks 2537 13.5 0.260 0.260 0.273 0.273 acc_transpose_blocks_sync 12462 16.5 0.262 0.266 0.262 0.266 dbcsr_create_new 23123 14.8 0.124 0.125 0.262 0.263 dbcsr_data_copy_aa2 2361 15.5 0.261 0.263 0.261 0.263 acc_transpose_blocks_kernels 4154 16.5 0.060 0.060 0.258 0.258 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="w64PBE", label="w64PBE", y=258.118, yerr=0.0 Plot: name="w64PBE_timings_6cpu_1gpu", title="Timings of w64PBE with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="rest", label="rest", y=137.659, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=42.324, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=33.589, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=19.983, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="pbe_lda_eval", label="pbe_lda_eval", y=12.99, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=11.573, yerr=0.0 Running w64SCAN.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/w64SCAN_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.273 0.273 989.263 989.263 qs_mol_dyn_low 1 2.0 0.005 0.005 986.637 986.640 qs_forces 11 3.9 0.003 0.003 986.582 986.582 qs_energies 11 4.9 0.002 0.002 888.118 888.119 scf_env_do_scf 11 5.9 0.001 0.002 846.013 846.014 velocity_verlet 10 3.0 0.002 0.002 785.665 785.685 scf_env_do_scf_inner_loop 106 6.8 0.008 0.011 759.237 759.237 rebuild_ks_matrix 117 8.5 0.001 0.001 680.196 680.196 qs_ks_build_kohn_sham_matrix 117 9.5 0.033 0.033 680.195 680.195 qs_ks_update_qs_env 119 7.8 0.002 0.002 598.315 598.317 fft_wrap_pw1pw2 3053 12.6 0.086 0.088 474.274 474.491 fft_wrap_pw1pw2_400 1649 13.9 0.013 0.013 454.584 454.702 qs_vxc_create 117 10.5 0.003 0.003 422.178 422.211 xc_vxc_pw_create 117 11.5 5.301 5.344 422.175 422.208 xc_rho_set_and_dset_create 117 12.5 6.771 6.827 279.120 279.842 qs_rho_update_rho_low 117 7.9 0.002 0.002 243.791 243.796 calculate_rho_elec 234 8.9 7.387 7.403 243.789 243.794 pw_gpu_c1dr3d_3d_ps 1521 15.1 132.387 133.073 237.785 237.793 pw_gpu_r3dc1d_3d_ps 1532 14.1 134.361 134.908 236.379 236.586 xc_pw_derive 702 13.5 0.016 0.016 206.302 206.806 sum_up_and_integrate 117 10.5 0.009 0.009 197.055 197.493 integrate_v_rspace 234 11.5 0.472 0.473 195.964 196.406 density_rs2pw 234 9.9 0.026 0.026 181.362 181.796 xc_functional_eval 234 13.5 0.004 0.005 164.780 165.483 libxc_spin_unpolarized_eval 234 14.5 164.169 164.869 164.776 165.478 xc_pw_divergence 117 12.5 0.008 0.009 136.085 136.593 potential_pw2rs 234 12.5 0.324 0.331 107.361 107.573 grid_integrate_task_list 234 12.5 88.129 88.781 88.129 88.781 init_scf_loop 13 6.8 0.001 0.001 86.722 86.722 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 82.722 82.722 mp_alltoall_z22v 3053 16.6 78.717 80.401 78.717 80.401 grid_collocate_task_list 234 9.9 54.867 55.254 54.867 55.254 x_to_yz 1521 16.1 12.083 12.088 50.478 51.272 yz_to_x 1532 15.1 9.447 9.466 49.768 50.645 transfer_rs2pw 947 10.9 0.023 0.024 39.764 40.111 transfer_rs2pw_400 245 11.8 28.002 28.067 34.498 34.848 pw_gpu_fg 1532 15.1 33.946 34.066 33.946 34.066 pw_gpu_sf 1521 16.1 32.849 32.863 32.849 32.863 transfer_pw2rs 947 13.5 0.019 0.020 32.644 32.662 transfer_pw2rs_400 245 14.3 23.023 23.110 28.937 28.957 init_scf_run 11 5.9 0.000 0.000 26.501 26.501 scf_env_initial_rho_setup 11 6.9 0.000 0.001 26.500 26.500 wfi_extrapolate 11 7.9 0.002 0.002 22.554 22.554 pw_gpu_ffc 1521 16.1 22.039 22.126 22.039 22.126 dbcsr_multiply_generic 2139 12.6 0.175 0.176 21.176 21.634 pw_poisson_solve 117 10.5 0.005 0.005 20.764 20.773 pw_gpu_cff 1532 15.1 18.136 18.137 18.136 18.137 pw_derive 1053 13.8 15.892 15.907 15.892 15.907 build_core_hamiltonian_matrix_ 11 4.9 0.002 0.002 15.605 15.676 fft_wrap_pw1pw2_140 468 13.2 0.004 0.004 15.444 15.516 qs_scf_new_mos 106 7.8 0.001 0.001 15.508 15.510 qs_scf_loop_do_ot 106 8.8 0.001 0.001 15.507 15.509 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 15.056 15.056 ot_scf_mini 106 9.8 0.004 0.004 13.954 13.955 pw_copy 2223 13.1 11.631 11.653 11.631 11.653 prepare_preconditioner 13 7.8 0.000 0.000 10.964 10.966 make_preconditioner 13 8.8 0.000 0.000 10.963 10.966 multiply_cannon 2139 13.6 0.394 0.396 10.208 10.225 mp_waitall_1 60839 17.0 9.839 9.840 9.839 9.840 pw_integral_ab_c1d_c1d_gs 117 11.5 8.836 8.838 9.216 9.228 multiply_cannon_loop 2139 14.6 0.294 0.296 9.005 9.012 make_m2s 4278 13.6 0.053 0.053 8.697 8.697 ot_mini 106 10.8 0.001 0.001 8.560 8.561 pw_poisson_set 118 11.5 0.009 0.009 8.539 8.548 make_images 4278 14.6 1.108 1.118 8.493 8.493 mp_sendrecv_dv 947 12.9 7.905 8.298 7.905 8.298 qs_env_update_s_mstruct 11 6.9 0.000 0.000 7.716 7.787 pw_axpy 1638 11.7 7.638 7.647 7.638 7.647 build_core_ppl_forces 11 5.9 6.990 7.062 6.990 7.062 build_core_hamiltonian_matrix 11 6.9 0.002 0.002 6.917 6.940 make_full_inverse_cholesky 13 9.8 0.000 0.000 5.821 5.975 build_kinetic_matrix_low 22 6.9 5.765 5.773 5.858 5.866 calculate_rho_core 11 7.9 0.487 0.489 5.466 5.491 qs_ot_get_derivative 106 11.8 0.002 0.002 5.448 5.451 build_overlap_matrix_low 22 6.9 5.034 5.059 5.126 5.152 multiply_cannon_multrec 4278 15.6 1.856 1.860 4.693 4.707 make_full_single_inverse 13 9.8 0.002 0.002 4.505 4.505 make_images_data 4278 15.6 0.066 0.067 4.474 4.482 transfer_rs2pw_140 234 11.9 3.326 3.345 4.466 4.480 hybrid_alltoall_any 4278 16.6 3.119 3.127 4.209 4.215 copy_dbcsr_to_fm 138 10.8 0.094 0.096 3.655 3.659 transfer_dbcsr_to_fm 13 10.8 0.004 0.006 3.254 3.254 fft_wrap_pw1pw2_50 468 13.2 0.003 0.003 3.169 3.187 ot_diis_step 106 11.8 0.006 0.006 3.091 3.091 dbcsr_to_fm_plan_create 13 12.8 2.692 2.692 3.038 3.050 transfer_pw2rs_140 234 14.5 1.904 1.907 2.991 2.994 build_core_ppl 11 7.9 2.866 2.877 2.866 2.877 arnoldi_generalized_ev 13 10.8 0.000 0.000 2.828 2.828 pw_zero 702 12.6 2.792 2.809 2.792 2.809 dbcsr_sym_matrix_vector_mult 1206 12.5 0.043 0.043 2.768 2.770 apply_preconditioner_dbcsr 119 12.8 0.000 0.000 2.629 2.633 apply_single 119 13.8 0.001 0.001 2.629 2.633 qs_ot_get_derivative_taylor 89 12.9 0.005 0.005 2.613 2.613 gev_build_subspace 22 11.5 0.014 0.014 2.603 2.603 dbcsr_mm_accdrv_process 9536 16.3 1.049 1.158 2.524 2.534 dbcsr_sym_matrix_vector_mult_l 1206 13.5 2.381 2.385 2.389 2.392 calculate_dm_sparse 117 9.7 0.001 0.002 2.339 2.343 qs_init_subsys 1 2.0 0.001 0.001 2.254 2.254 qs_env_setup 1 3.0 0.000 0.000 2.245 2.246 qs_env_rebuild_pw_env 23 5.3 0.000 0.000 2.244 2.246 pw_env_rebuild 1 5.0 0.000 0.000 2.244 2.245 pw_grid_setup 4 6.0 0.000 0.000 2.172 2.173 pw_grid_setup_internal 4 7.0 0.022 0.022 2.137 2.139 cp_dbcsr_sm_fm_multiply 45 9.4 0.002 0.002 2.099 2.101 qs_create_task_list 11 7.9 0.000 0.000 2.013 2.064 generate_qs_task_list 11 8.9 1.002 1.002 2.013 2.063 multiply_cannon_sync_h2d 4278 15.6 1.925 1.978 1.925 1.978 qs_ot_get_p 119 10.6 0.002 0.002 1.953 1.956 dbcsr_special_finalize 6417 15.6 0.041 0.042 1.836 1.837 pw_grid_sort 4 8.0 1.277 1.305 1.737 1.774 dbcsr_merge_single_wm 4278 16.6 0.156 0.158 1.702 1.704 cp_dbcsr_sm_fm_multiply_core 45 10.4 0.000 0.000 1.635 1.638 dbcsr_complete_redistribute 299 11.7 0.597 0.606 1.471 1.620 integrate_v_core_rspace 11 7.9 0.177 0.178 1.486 1.487 multiply_cannon_metrocomm1 4278 15.6 0.015 0.015 1.424 1.478 pw_scale 585 11.9 1.460 1.470 1.460 1.470 cp_fm_cholesky_invert 13 10.8 1.449 1.449 1.449 1.449 mp_sum_d 3883 11.6 1.216 1.427 1.216 1.427 copy_fm_to_dbcsr 174 10.8 0.002 0.002 1.255 1.412 dbcsr_dot 1134 12.2 1.259 1.263 1.341 1.350 mp_sum_l 6446 13.6 0.915 1.343 0.915 1.343 calculate_first_density_matrix 1 7.0 0.000 0.000 1.314 1.314 dbcsr_sort_data 4278 17.6 1.239 1.242 1.239 1.242 cp_dbcsr_plus_fm_fm_t 22 8.9 0.001 0.001 1.055 1.056 jit_kernel_multiply 10 15.0 0.915 1.012 0.915 1.012 pw_multiply_with 117 11.5 0.997 1.003 0.997 1.003 fft_wrap_pw1pw2_20 468 13.2 0.003 0.003 0.991 0.999 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="w64SCAN", label="w64SCAN", y=989.263, yerr=0.0 Plot: name="w64SCAN_timings_6cpu_1gpu", title="Timings of w64SCAN with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="rest", label="rest", y=391.5, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="libxc_spin_unpolarized_eval", label="libxc_spin_unpolarized_eval", y=164.169, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="pw_gpu_r3dc1d_3d_ps", label="pw_gpu_r3dc1d_3d_ps", y=134.361, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=132.387, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=88.129, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=78.717, yerr=0.0 Running ZnO.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/ZnO_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.028 0.030 97.340 97.340 qs_energies 1 2.0 0.001 0.001 95.866 95.866 scf_env_do_scf 1 3.0 0.000 0.000 94.774 94.774 scf_env_do_scf_inner_loop 10 4.0 0.005 0.005 94.774 94.774 qs_scf_new_mos_kp 10 5.0 0.000 0.000 92.152 92.229 do_general_diag_kp 10 6.0 37.496 37.612 92.152 92.229 cp_cfm_geeig_local 15051 7.0 33.268 33.295 33.268 33.295 kpoint_density_transform 10 7.0 0.016 0.016 19.809 19.814 kpoint_density_transform_regul 10 8.0 0.168 0.172 19.784 19.784 kp_density_fft 10 9.0 2.552 2.563 15.892 15.892 k_grid_to_cell_fft 910 10.0 3.724 3.790 10.953 11.055 fft3d_s 28881 11.0 7.201 7.239 7.230 7.265 mp_alltoall_z11v 910 10.0 2.387 2.499 2.387 2.499 qs_ks_update_qs_env 10 5.0 0.000 0.000 1.920 1.997 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="ZnO", label="ZnO", y=97.34, yerr=0.0 Plot: name="ZnO_timings_6cpu_1gpu", title="Timings of ZnO with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="rest", label="rest", y=13.099000000000004, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="do_general_diag_kp", label="do_general_diag_kp", y=37.496, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="cp_cfm_geeig_local", label="cp_cfm_geeig_local", y=33.268, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="fft3d_s", label="fft3d_s", y=7.201, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="k_grid_to_cell_fft", label="k_grid_to_cell_fft", y=3.724, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="kp_density_fft", label="kp_density_fft", y=2.552, yerr=0.0 Running GW_PBE_4benzene.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/GW_PBE_4benzene_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.024 0.025 116.482 116.482 qs_energies 1 2.0 0.000 0.000 116.102 116.106 mp2_main 1 3.0 0.000 0.000 108.368 108.371 mp2_gpw_main 1 4.0 0.000 0.000 106.461 106.465 rpa_ri_compute_en 1 5.0 0.000 0.000 97.744 97.748 rpa_num_int 1 6.0 0.001 0.001 97.735 97.738 dbt_total 2336 9.6 0.025 0.025 77.843 77.843 compute_mat_P_omega 1 7.0 0.002 0.003 74.131 74.153 compute_mat_P_omega_contract 10 8.0 5.848 5.885 73.755 73.764 dbt_contract 787 11.0 0.054 0.054 50.071 50.073 dbt_tas_total 1149 12.2 0.167 0.168 38.425 38.425 dbt_tas_multiply 807 12.1 0.004 0.004 37.663 37.664 dbt_copy 1107 10.7 0.074 0.075 28.148 28.284 dbt_tas_dbm 807 14.1 0.007 0.007 28.274 28.274 dbm_multiply 807 16.1 27.355 27.748 27.355 27.748 compute_mat_P_omega_calc_M_occ 250 9.0 5.860 5.912 25.615 25.615 dbt_reshape 594 11.8 8.016 8.273 18.832 19.059 dbt_tas_mm_1N 524 15.1 0.003 0.003 18.265 18.575 compute_QP_energies 1 7.0 0.000 0.000 16.953 16.953 compute_self_energy_cubic_gw 1 8.0 0.137 0.139 16.952 16.952 compute_mat_P_omega_calc_M_vir 250 9.0 0.001 0.002 15.708 15.708 dbt_tas_reserve_blocks_index 3266 14.3 0.719 0.728 11.462 11.627 dbm_reserve_blocks 3634 15.3 11.072 11.235 11.072 11.235 dbt_crop 1042 12.0 7.390 7.466 9.780 9.883 dbt_reserve_blocks_index 2347 13.0 0.341 0.343 9.747 9.820 dbt_reserve_blocks_index_array 2289 12.1 0.013 0.013 9.533 9.625 compute_mat_P_omega_calc_P_t 250 9.0 0.001 0.001 9.604 9.604 mp_waitall_2 2656 15.9 8.907 8.920 8.907 8.920 mp2_ri_gpw_compute_in 1 5.0 0.001 0.001 8.706 8.706 dbt_tas_mm_2 251 15.0 0.003 0.003 8.073 8.073 dbt_communicate_buffer 594 12.8 0.016 0.016 8.008 8.040 contract_cubic_gw 21 9.0 0.000 0.000 7.782 7.782 scf_env_do_scf 1 3.0 0.000 0.000 7.099 7.099 scf_env_do_scf_inner_loop 17 4.0 0.001 0.002 7.099 7.099 compute_mat_P_omega_copy_M_vir 250 9.0 0.002 0.002 6.041 6.065 compute_mat_P_omega_copy_M_occ 250 9.0 0.002 0.002 5.857 5.868 dbcsr_multiply_generic 30 8.1 0.003 0.003 5.033 5.109 dbt_tas_copy 511 11.5 2.856 2.889 4.705 4.971 multiply_cannon 30 9.1 0.007 0.008 4.814 4.886 multiply_cannon_loop 30 10.1 0.005 0.005 4.753 4.825 multiply_cannon_multrec 60 11.1 0.270 0.281 4.132 4.180 get_2c_integrals 1 6.0 0.000 0.000 3.908 3.909 qs_scf_new_mos 17 5.0 0.001 0.001 3.852 3.901 trace_sigma_gw 21 9.0 0.577 0.586 3.733 3.733 dbcsr_mm_accdrv_process 328 12.3 0.026 0.027 3.554 3.579 jit_kernel_multiply 17 11.6 3.520 3.545 3.520 3.545 dbt_split_copyback 70 10.6 1.291 1.427 3.007 3.166 mp_sync 8688 11.6 2.775 3.079 2.775 3.079 compute_2c_integrals 1 7.0 0.000 0.000 3.071 3.071 fft_wrap_pw1pw2 301 10.2 0.006 0.007 2.776 2.782 convert_to_new_pgrid 2421 14.1 0.043 0.043 2.770 2.772 dbm_copy 1614 15.1 2.727 2.729 2.727 2.729 qs_ks_build_kohn_sham_matrix 18 6.9 0.004 0.004 2.663 2.663 mp2_ri_gpw_compute_in_copy_3c 6 6.0 0.246 0.252 2.479 2.650 qs_ks_update_qs_env 17 5.0 0.000 0.000 2.628 2.629 rebuild_ks_matrix 17 6.0 0.000 0.000 2.621 2.621 fill_fm_L_from_L_loc_non_block 1 8.0 0.000 0.000 2.487 2.496 build_3c_integrals 5 6.0 1.623 1.645 2.309 2.480 fill_fm_L_from_L_loc_non_block 1 9.0 2.382 2.391 2.382 2.391 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="GW_PBE_4benzene", label="GW_PBE_4benzene", y=116.482, yerr=0.0 Plot: name="GW_PBE_4benzene_timings_6cpu_1gpu", title="Timings of GW_PBE_4benzene with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="rest", label="rest", y=53.742, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbm_multiply", label="dbm_multiply", y=27.355, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbm_reserve_blocks", label="dbm_reserve_blocks", y=11.072, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="mp_waitall_2", label="mp_waitall_2", y=8.907, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbt_reshape", label="dbt_reshape", y=8.016, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbt_crop", label="dbt_crop", y=7.39, yerr=0.0 Running RI-HFX_H2O-32.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/RI-HFX_H2O-32_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.026 0.027 230.050 230.050 qs_forces 1 2.0 0.000 0.000 229.529 229.529 rebuild_ks_matrix 7 6.6 0.000 0.000 224.633 224.634 qs_ks_build_kohn_sham_matrix 7 7.6 0.003 0.003 224.633 224.634 hfx_ks_matrix 7 8.6 0.000 0.000 220.322 220.333 dbt_total 849 11.0 0.011 0.011 164.027 164.027 hfx_ri_update_ks 7 9.6 0.000 0.000 128.443 128.443 hfx_ri_update_ks_Pmat 7 10.6 26.093 26.163 128.438 128.438 qs_energies 1 3.0 0.000 0.000 123.861 123.862 scf_env_do_scf 1 4.0 0.000 0.000 121.626 121.626 qs_ks_update_qs_env 8 6.0 0.000 0.000 119.021 119.022 qs_ks_update_qs_env_forces 1 3.0 0.000 0.000 105.620 105.620 dbt_contract 207 12.4 0.061 0.061 94.702 94.702 hfx_ri_update_forces 1 7.0 1.217 1.246 91.876 91.888 dbt_tas_total 369 13.4 0.097 0.098 76.985 76.985 dbt_tas_multiply 216 13.5 0.001 0.001 73.717 73.717 scf_env_do_scf_inner_loop 6 5.0 0.001 0.001 64.668 64.668 dbt_copy 423 11.8 0.050 0.050 63.478 63.952 dbt_tas_dbm 216 15.5 0.002 0.002 57.821 57.821 init_scf_loop 2 5.0 0.000 0.000 56.956 56.956 hfx_ri_forces_Pmat_3c 1 8.0 4.083 4.087 54.926 54.934 dbm_multiply 216 17.5 53.870 53.944 53.870 53.944 dbt_reshape 175 13.2 21.765 22.248 48.598 49.092 hfx_ri_update_ks_Pmat_KS 63 11.6 0.001 0.001 36.623 36.623 precalc_derivatives 1 8.0 2.051 2.071 30.388 30.388 mp_waitall_2 1022 16.5 24.948 24.997 24.948 24.997 dbt_tas_mm_2 91 16.5 0.001 0.001 24.370 24.370 dbt_communicate_buffer 175 14.2 0.006 0.006 20.665 20.753 dbt_crop 372 13.7 16.037 16.277 20.381 20.626 dbt_tas_reserve_blocks_index 1323 15.4 1.861 1.888 19.673 19.682 dbt_tas_mm_3T 77 17.1 0.001 0.001 18.687 18.964 hfx_ri_update_ks_Pmat_copy_2 63 11.6 0.000 0.000 18.932 18.932 dbm_reserve_blocks 1491 16.3 18.505 18.537 18.505 18.537 hfx_ri_pre_scf_Pmat 1 12.0 0.000 0.000 18.335 18.335 hfx_ri_update_ks_Pmat_Px3C 63 11.6 0.000 0.000 17.783 17.783 build_3c_derivatives 3 9.0 2.725 2.794 16.706 16.708 dbt_reserve_blocks_index 889 14.5 0.684 0.706 15.774 15.886 dbt_reserve_blocks_index_array 859 13.5 0.010 0.010 15.466 15.575 dbt_tas_mm_3N 37 15.4 0.000 0.000 11.784 11.796 dbt_tas_copy 248 12.5 5.033 5.036 9.356 9.461 mp_sync 2901 12.8 8.588 9.120 8.588 9.120 hfx_ri_pre_scf_Pmat_int 1 13.0 0.000 0.000 5.874 5.874 dbt_tas_replicate 168 15.1 2.570 2.611 5.661 5.674 hfx_ri_pre_scf_calc_tensors 1 14.0 0.004 0.004 5.080 5.080 hfx_ri_pre_scf_Pmat_copy_2 9 13.0 1.858 1.865 4.883 4.889 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="RI-HFX_H2O-32", label="RI-HFX_H2O-32", y=230.05, yerr=0.0 Plot: name="RI-HFX_H2O-32_timings_6cpu_1gpu", title="Timings of RI-HFX_H2O-32 with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="rest", label="rest", y=84.86900000000003, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbm_multiply", label="dbm_multiply", y=53.87, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="hfx_ri_update_ks_Pmat", label="hfx_ri_update_ks_Pmat", y=26.093, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="mp_waitall_2", label="mp_waitall_2", y=24.948, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbt_reshape", label="dbt_reshape", y=21.765, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbm_reserve_blocks", label="dbm_reserve_blocks", y=18.505, yerr=0.0 Running RI-MP2_ammonia.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/RI-MP2_ammonia_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.014 0.016 118.416 118.416 qs_energies 1 2.0 0.000 0.000 118.190 118.190 mp2_main 1 3.0 0.001 0.001 108.646 108.646 mp2_gpw_main 1 4.0 0.002 0.002 107.978 107.978 mp2_ri_gpw_compute_in 1 5.0 0.610 0.611 57.822 57.825 mp2_ri_gpw_compute_en 1 5.0 0.114 0.116 50.081 50.083 mp2_ri_gpw_compute_in_loop 1 6.0 0.019 0.021 48.743 48.747 mp2_ri_gpw_compute_en_RI_loop 1 6.0 13.353 13.390 47.129 47.134 dbcsr_multiply_generic 2666 8.0 0.208 0.210 26.990 27.084 ao_to_mo_and_store_B_mult_1 1328 7.0 0.020 0.020 25.919 26.012 mp2_ri_gpw_compute_en_expansio 1040 7.0 0.998 1.024 18.980 19.221 local_gemm 1040 8.0 17.982 18.198 17.982 18.198 mp2_eri_3c_integrate_gpw 1328 7.0 0.024 0.025 16.953 17.008 make_m2s 5332 9.0 0.075 0.077 16.230 16.278 make_images 5332 10.0 2.444 2.447 16.012 16.061 make_images_data 5332 11.0 0.090 0.090 11.536 11.554 hybrid_alltoall_any 5332 12.0 11.271 11.284 11.310 11.325 multiply_cannon 2666 9.0 0.489 0.489 9.940 10.090 multiply_cannon_loop 2666 10.0 0.235 0.236 8.595 8.744 scf_env_do_scf 1 3.0 0.000 0.000 8.507 8.508 scf_env_do_scf_inner_loop 10 4.0 0.001 0.001 8.507 8.508 get_2c_integrals 1 6.0 0.005 0.005 8.463 8.468 integrate_v_rspace 1338 8.0 1.178 1.188 8.134 8.143 fft_wrap_pw1pw2 26668 10.4 0.166 0.173 8.081 8.133 compute_2c_integrals 1 7.0 0.008 0.008 7.797 7.798 compute_2c_integrals_loop_lm 1 8.0 0.024 0.025 7.573 7.610 mp2_eri_2c_integrate_gpw 1 9.0 2.267 2.282 7.549 7.586 collocate_function 1328 8.0 5.490 5.607 7.513 7.564 mp2_ri_gpw_compute_en_comm 221 7.0 1.272 1.276 7.039 7.440 qs_scf_new_mos 10 5.0 0.000 0.000 6.698 6.702 mp2_ri_gpw_compute_en_ener 1040 7.0 6.473 6.583 6.473 6.583 ao_to_mo_and_store_B_E_Ex_1 1328 7.0 3.947 3.954 5.586 5.619 grid_integrate_task_list 1338 9.0 5.451 5.461 5.451 5.461 mp_sendrecv_dm3 442 8.0 4.438 4.849 4.438 4.849 fft_wrap_pw1pw2_20 10647 11.4 0.026 0.027 4.586 4.627 multiply_cannon_multrec 2676 11.0 1.869 1.969 4.258 4.328 pw_gpu_r3dc1d_3d 13282 12.2 3.834 3.877 3.834 3.877 eigensolver 11 5.8 0.002 0.002 3.816 3.821 copy_dbcsr_to_fm 1351 8.0 0.106 0.106 3.745 3.767 cp_fm_diag_elpa 11 6.8 0.000 0.000 3.004 3.005 cp_fm_diag_elpa_base 11 7.8 2.900 2.924 3.003 3.003 potential_pw2rs 2666 10.0 0.121 0.123 2.972 2.995 pw_gpu_c1dr3d_3d 13280 12.7 2.864 2.877 2.864 2.877 fill_local_i_aL 884 7.5 2.613 2.627 2.613 2.627 replicate_iaK_2intgroup 1 6.0 2.354 2.375 2.503 2.526 collocate_single_gaussian 1328 10.0 0.117 0.118 2.467 2.473 mp2_eri_2c_integrate_gpw_pot_l 1328 10.0 0.005 0.006 2.452 2.471 fft_wrap_pw1pw2_10 15957 11.5 0.024 0.024 2.438 2.455 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="RI-MP2_ammonia", label="RI-MP2_ammonia", y=118.416, yerr=0.0 Plot: name="RI-MP2_ammonia_timings_6cpu_1gpu", title="Timings of RI-MP2_ammonia with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="rest", label="rest", y=63.846999999999994, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="local_gemm", label="local_gemm", y=17.982, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="mp2_ri_gpw_compute_en_RI_loop", label="mp2_ri_gpw_compute_en_RI_loop", y=13.353, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=11.271, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="mp2_ri_gpw_compute_en_ener", label="mp2_ri_gpw_compute_en_ener", y=6.473, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="collocate_function", label="collocate_function", y=5.49, yerr=0.0 Running diag_cu144_broy.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/diag_cu144_broy_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.117 0.120 212.693 212.693 qs_energies 1 2.0 0.000 0.000 211.442 211.442 scf_env_do_scf 1 3.0 0.000 0.000 196.354 196.354 scf_env_do_scf_inner_loop 15 4.0 0.002 0.002 196.354 196.354 qs_ks_update_qs_env 15 5.0 0.000 0.000 91.385 91.424 rebuild_ks_matrix 15 6.0 0.000 0.000 91.162 91.201 qs_ks_build_kohn_sham_matrix 15 7.0 0.004 0.004 91.162 91.201 qs_scf_new_mos 15 5.0 0.130 0.131 71.389 71.431 fft_wrap_pw1pw2 1071 10.6 0.033 0.034 59.885 59.908 eigensolver 15 6.0 0.003 0.003 52.156 52.236 qs_vxc_create 15 8.0 0.195 0.197 46.657 46.662 sum_up_and_integrate 15 8.0 0.001 0.001 42.696 42.741 integrate_v_rspace 15 9.0 0.055 0.056 42.662 42.707 grid_integrate_task_list 15 10.0 34.378 34.418 34.378 34.418 cp_fm_diag_elpa 15 7.0 0.000 0.000 31.330 31.335 cp_fm_diag_elpa_base 15 8.0 29.449 30.024 31.323 31.324 qs_rho_update_rho_low 16 5.0 0.000 0.000 30.696 30.696 calculate_rho_elec 16 6.0 0.205 0.206 30.696 30.696 calculate_vxc_nlvdw 15 9.0 1.169 1.172 30.374 30.379 fft_wrap_pw1pw2_150 735 11.9 0.005 0.006 30.283 30.284 pw_gpu_c1dr3d_3d_ps 555 12.7 6.255 6.289 29.990 30.008 pw_gpu_r3dc1d_3d_ps 516 12.6 6.958 7.050 29.854 29.859 cp_fm_cholesky_restore 45 7.0 18.809 19.567 18.809 19.567 grid_collocate_task_list 16 7.0 17.498 17.517 17.498 17.517 copy_dbcsr_to_fm 16 5.9 0.556 0.577 16.186 16.243 fft_wrap_pw1pw2_200 212 11.3 0.002 0.002 15.999 16.069 dbcsr_to_fm_plan_create 16 6.9 13.764 13.888 15.207 15.247 vdW_theta_forward 15 10.0 0.719 0.722 13.523 13.530 density_rs2pw 16 7.0 0.002 0.002 12.973 12.997 vdW_theta_inverse 15 10.0 0.495 0.501 10.997 10.998 qs_energies_init_hamiltonians 1 3.0 0.000 0.000 10.845 10.845 mp_alltoall_z22v 1071 14.6 10.037 10.159 10.037 10.159 pw_gpu_ffc 555 13.7 9.782 9.806 9.782 9.806 pw_gpu_cff 516 13.6 9.423 9.452 9.423 9.452 build_core_hamiltonian_matrix 1 4.0 0.000 0.000 9.406 9.431 xc_vxc_pw_create 15 9.0 0.206 0.207 8.687 8.687 potential_pw2rs 15 10.0 0.008 0.008 8.229 8.235 pw_gpu_fg 516 13.6 7.511 7.525 7.511 7.525 pw_gpu_sf 555 13.7 7.242 7.252 7.242 7.252 x_to_yz 555 13.7 1.401 1.403 6.674 6.725 fft_wrap_pw1pw2_10 62 10.5 0.000 0.000 6.173 6.176 yz_to_x 516 13.6 1.135 1.137 5.898 5.970 xc_pw_derive 90 11.0 0.002 0.002 5.582 5.588 prepare_nlvdw_density 15 9.0 0.074 0.075 5.535 5.541 build_core_ppnl 1 5.0 5.267 5.294 5.267 5.294 cp_fm_uplo_to_full 30 8.0 3.889 5.136 3.889 5.136 xc_rho_set_and_dset_create 15 10.0 0.143 0.144 4.702 4.703 gspace_mixing 14 5.0 0.139 0.139 4.597 4.597 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="diag_cu144_broy", label="diag_cu144_broy", y=212.693, yerr=0.0 Plot: name="diag_cu144_broy_timings_6cpu_1gpu", title="Timings of diag_cu144_broy with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="rest", label="rest", y=98.79500000000002, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=34.378, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="cp_fm_diag_elpa_base", label="cp_fm_diag_elpa_base", y=29.449, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="cp_fm_cholesky_restore", label="cp_fm_cholesky_restore", y=18.809, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=17.498, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=13.764, yerr=0.0 Running bench_dftb.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/bench_dftb_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 2.382 2.399 178.619 178.619 qs_energies 1 2.0 0.000 0.000 176.117 176.118 ls_scf 1 3.0 0.000 0.000 168.060 168.062 ls_scf_main 1 4.0 0.001 0.001 155.523 155.524 density_matrix_trs4 5 5.0 0.005 0.005 122.908 122.947 dbcsr_multiply_generic 95 6.2 0.198 0.216 106.205 106.265 multiply_cannon 95 7.2 2.237 2.566 75.274 75.437 multiply_cannon_loop 95 8.2 0.186 0.188 63.146 63.255 multiply_cannon_multrec 190 9.2 48.184 48.285 54.096 54.209 ls_scf_dm_to_ks 5 5.0 0.000 0.000 30.497 30.539 make_m2s 190 7.2 0.018 0.018 26.062 26.101 make_images 190 8.2 5.827 5.930 25.456 25.494 matrix_ls_to_qs 5 6.0 0.000 0.000 19.589 19.705 dbcsr_complete_redistribute 11 7.5 11.861 11.885 16.845 16.984 matrix_decluster 5 7.0 0.000 0.000 15.339 15.491 qs_ks_update_qs_env 6 6.2 0.000 0.000 13.074 13.150 arnoldi_extremal 6 6.2 0.000 0.000 12.855 12.857 arnoldi_normal_ev 6 7.2 0.006 0.006 12.854 12.857 rebuild_ks_matrix 6 7.2 0.000 0.000 12.599 12.606 build_dftb_ks_matrix 6 8.2 0.001 0.001 12.599 12.606 build_subspace 12 8.2 0.037 0.038 12.560 12.561 build_dftb_coulomb 6 9.2 0.968 1.042 12.270 12.277 dbcsr_matrix_vector_mult 310 9.0 0.086 0.087 11.291 11.575 dbcsr_matrix_vector_mult_local 310 10.0 10.713 11.000 10.718 11.005 tb_ewald_overlap 6 10.2 10.883 11.004 10.883 11.004 ls_scf_init_scf 1 4.0 0.000 0.000 10.645 10.649 make_images_data 190 9.2 0.008 0.009 10.415 10.534 hybrid_alltoall_any 201 10.0 6.871 6.892 9.951 10.076 dbcsr_finalize 277 7.6 0.111 0.112 8.447 8.512 calculate_norms 380 9.2 8.475 8.500 8.475 8.500 ls_scf_init_matrix_S 1 5.0 0.000 0.000 8.443 8.445 qs_energies_init_hamiltonians 1 3.0 0.000 0.000 7.988 7.989 dbcsr_merge_all 247 8.6 1.642 1.665 7.755 7.809 matrix_sqrt_Newton_Schulz 1 6.0 0.001 0.001 7.667 7.670 build_qs_neighbor_lists 1 4.0 0.000 0.000 7.264 7.334 build_neighbor_lists_sab_tbe 1 5.0 7.057 7.123 7.057 7.123 dbcsr_copy 443 8.0 1.108 1.119 5.373 5.398 dbcsr_special_finalize 285 9.2 0.007 0.007 5.387 5.396 setup_rec_index_2d 190 8.2 5.271 5.307 5.271 5.307 dbcsr_sort_indices 643 10.1 5.044 5.049 5.044 5.049 dbcsr_data_new 3509 9.3 4.568 4.897 4.568 4.897 dbcsr_mm_accdrv_process 8119 10.0 0.540 0.601 4.746 4.753 dbcsr_add_d 130 6.0 0.001 0.001 4.708 4.730 dbcsr_add_anytype 130 7.0 1.986 1.989 4.707 4.729 dbcsr_dot 66 6.3 4.188 4.189 4.555 4.692 dbcsr_copy_into_existing 5 8.0 4.249 4.286 4.249 4.286 dbcsr_mm_accdrv_process_sort 8119 11.0 4.144 4.152 4.144 4.152 tree_to_linear_d 11 10.5 3.950 3.958 3.950 3.958 dbcsr_mm_multrec_init 95 8.2 0.000 0.000 3.573 3.895 dbcsr_mm_csr_init 95 9.2 0.006 0.006 3.572 3.895 dbcsr_mm_sched_init 95 10.2 0.000 0.000 3.541 3.863 dbcsr_mm_accdrv_init 95 11.2 0.313 0.334 3.541 3.863 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="bench_dftb", label="bench_dftb", y=178.619, yerr=0.0 Plot: name="bench_dftb_timings_6cpu_1gpu", title="Timings of bench_dftb with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="rest", label="rest", y=88.503, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="multiply_cannon_multrec", label="multiply_cannon_multrec", y=48.184, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="dbcsr_complete_redistribute", label="dbcsr_complete_redistribute", y=11.861, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="tb_ewald_overlap", label="tb_ewald_overlap", y=10.883, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="dbcsr_matrix_vector_mult_local", label="dbcsr_matrix_vector_mult_local", y=10.713, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="calculate_norms", label="calculate_norms", y=8.475, yerr=0.0 Running dbcsr.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/dbcsr_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.006 0.007 54.042 54.042 lib_test 1 2.0 0.000 0.000 54.026 54.034 dbcsr_run_tests 3 3.0 0.001 0.001 54.025 54.032 test_multiplies_multiproc 3 4.0 0.001 0.001 41.750 41.821 dbcsr_multiply_generic 9 5.0 0.002 0.002 32.200 32.204 multiply_cannon 9 6.0 0.210 0.396 21.338 21.888 multiply_cannon_loop 9 7.0 0.003 0.003 19.765 20.210 multiply_cannon_multrec 18 8.0 10.312 10.701 18.368 18.835 dbcsr_make_random_matrix 9 4.0 8.381 8.577 12.117 12.191 dbcsr_finalize 27 5.7 0.001 0.001 8.176 8.273 dbcsr_merge_all 18 6.5 3.934 3.956 8.053 8.152 dbcsr_mm_accdrv_process 8199 9.0 1.475 1.589 7.822 7.903 dbcsr_redistribute 9 5.0 3.947 4.023 6.648 6.691 make_m2s 18 6.0 0.001 0.001 5.552 5.580 make_images 18 7.0 0.374 0.380 5.515 5.543 dbcsr_mm_accdrv_process_sort 8199 10.0 5.291 5.337 5.291 5.337 make_images_data 18 8.0 0.001 0.001 3.247 3.274 hybrid_alltoall_any 18 9.0 2.654 2.675 3.192 3.218 mp_alltoall_d11v 27 6.0 2.348 2.392 2.348 2.392 tree_to_linear_d 9 7.0 2.050 2.083 2.050 2.083 dbcsr_data_copy_aa2 18 7.5 1.917 2.072 1.917 2.072 dbcsr_data_release 507 7.7 1.558 1.569 1.558 1.569 dbcsr_data_new 354 7.4 1.084 1.212 1.084 1.212 mp_sum_l 61 4.9 0.607 1.212 0.607 1.212 dbcsr_multiply_generic_mpsum_f 9 6.0 0.000 0.000 0.607 1.211 jit_kernel_multiply 5 10.0 1.055 1.204 1.055 1.204 dbcsr_checksum 6 5.0 1.164 1.166 1.171 1.171 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="dbcsr", label="dbcsr", y=54.042, yerr=0.0 Plot: name="dbcsr_timings_6cpu_1gpu", title="Timings of dbcsr with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="rest", label="rest", y=22.177, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="multiply_cannon_multrec", label="multiply_cannon_multrec", y=10.312, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_make_random_matrix", label="dbcsr_make_random_matrix", y=8.381, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_mm_accdrv_process_sort", label="dbcsr_mm_accdrv_process_sort", y=5.291, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_redistribute", label="dbcsr_redistribute", y=3.947, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_merge_all", label="dbcsr_merge_all", y=3.934, yerr=0.0 Running MQAE_single_node.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/MQAE_single_node_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.051 0.052 231.974 231.974 qs_mol_dyn_low 1 2.0 0.005 0.006 230.137 230.175 qs_forces 6 3.8 0.001 0.001 143.684 143.684 qs_energies 6 4.8 0.001 0.001 135.512 135.512 scf_env_do_scf 6 5.8 0.001 0.001 127.832 127.832 scf_env_do_scf_inner_loop 113 6.2 0.007 0.010 119.712 119.712 velocity_verlet 5 3.0 0.004 0.004 111.426 111.480 rebuild_ks_matrix 119 8.1 0.001 0.001 99.249 99.249 qs_ks_build_kohn_sham_matrix 119 9.1 0.030 0.030 99.249 99.249 qs_ks_update_qs_env 119 7.3 0.002 0.002 93.721 93.722 fft_wrap_pw1pw2 2059 12.4 0.051 0.052 77.360 77.475 fft_wrap_pw1pw2_150 1321 13.9 0.012 0.012 74.163 74.261 qs_vxc_create 119 10.1 0.002 0.002 62.986 62.987 xc_vxc_pw_create 119 11.1 1.734 1.757 62.984 62.985 qmmm_el_coupling 6 3.8 0.000 0.000 46.442 46.471 qmmm_elec_with_gaussian 6 4.8 0.044 0.044 46.435 46.463 qmmm_elec_with_gaussian_low 6 5.8 0.000 0.000 43.993 44.614 xc_pw_derive 714 13.1 0.013 0.014 43.990 44.054 pw_gpu_c1dr3d_3d_ps 1095 14.8 11.672 11.814 41.556 41.591 qmmm_elec_gaussian_low_G 6 6.8 38.781 39.393 38.781 39.393 qmmm_forces 6 3.8 0.002 0.002 36.966 36.966 qmmm_forces_with_gaussian 6 4.8 0.054 0.054 36.346 36.543 pw_gpu_r3dc1d_3d_ps 964 14.0 10.646 10.784 35.736 35.818 qmmm_force_with_gaussian_low 6 5.8 0.000 0.000 34.716 34.908 xc_rho_set_and_dset_create 119 12.1 2.740 2.783 31.361 31.465 xc_pw_divergence 119 12.1 0.008 0.008 29.365 29.474 qmmm_forces_gaussian_low_G 6 6.8 29.167 29.382 29.167 29.382 qs_rho_update_rho_low 119 7.3 0.001 0.001 25.681 25.929 calculate_rho_elec 119 8.3 1.193 1.194 25.681 25.928 mp_alltoall_z22v 2059 16.4 19.577 20.085 19.577 20.085 density_rs2pw 119 9.3 0.010 0.011 19.255 19.505 sum_up_and_integrate 119 10.1 0.005 0.006 17.530 17.573 integrate_v_rspace 119 11.1 0.025 0.025 17.264 17.307 x_to_yz 1095 15.8 3.075 3.077 13.635 13.864 dbcsr_multiply_generic 2616 12.3 0.127 0.129 12.339 12.492 yz_to_x 964 15.0 2.257 2.260 11.274 11.553 potential_pw2rs 119 12.1 0.038 0.039 11.544 11.544 multiply_cannon 2616 13.3 0.280 0.283 10.406 10.699 qs_ks_ddapc 119 10.1 0.003 0.003 10.229 10.240 multiply_cannon_loop 2616 14.3 0.309 0.311 9.802 10.093 pw_gpu_sf 1095 15.8 8.996 9.015 8.996 9.015 pw_gpu_fg 964 15.0 8.645 8.721 8.645 8.721 init_scf_loop 6 6.8 0.000 0.000 8.117 8.117 qs_scf_new_mos 113 7.2 0.001 0.001 7.745 7.746 qs_scf_loop_do_ot 113 8.2 0.001 0.001 7.744 7.745 ot_scf_mini 113 9.2 0.003 0.003 7.423 7.424 multiply_cannon_multrec 5232 15.3 3.304 3.332 7.217 7.274 pw_gpu_ffc 1095 15.8 7.233 7.266 7.233 7.266 xc_functional_eval 238 13.1 0.004 0.004 5.735 5.790 pw_poisson_solve 125 9.9 0.004 0.005 5.752 5.775 grid_integrate_task_list 119 12.1 5.694 5.738 5.694 5.738 qmmm_forces_gaussian_low_R 6 6.8 0.000 0.000 5.549 5.571 qmmm_forces_with_gaussian_LG 6 7.8 5.549 5.571 5.549 5.571 qs_ks_update_qs_env_forces 6 4.8 0.000 0.000 5.563 5.563 pw_derive 1089 13.4 5.278 5.329 5.278 5.329 qmmm_elec_gaussian_low_R 6 6.8 0.000 0.000 5.212 5.221 qmmm_elec_with_gaussian_LG 6 7.8 5.212 5.221 5.212 5.221 grid_collocate_task_list 119 9.3 5.190 5.195 5.190 5.195 pw_gpu_cff 964 15.0 5.096 5.115 5.096 5.115 ot_mini 113 10.2 0.001 0.001 5.085 5.085 init_scf_run 6 5.8 0.000 0.000 5.056 5.056 scf_env_initial_rho_setup 6 6.8 0.000 0.000 5.056 5.056 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="MQAE_single_node", label="MQAE_single_node", y=231.974, yerr=0.0 Plot: name="MQAE_single_node_timings_6cpu_1gpu", title="Timings of MQAE_single_node with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="rest", label="rest", y=122.13099999999999, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="qmmm_elec_gaussian_low_G", label="qmmm_elec_gaussian_low_G", y=38.781, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="qmmm_forces_gaussian_low_G", label="qmmm_forces_gaussian_low_G", y=29.167, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=19.577, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=11.672, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="pw_gpu_r3dc1d_3d_ps", label="pw_gpu_r3dc1d_3d_ps", y=10.646, yerr=0.0 Summary: Performance test took 45 minutes. Status: OK ---> Removed intermediate container f1a5e1bc73cd ---> 19a3e5a2d69e Step 46/47 : CMD cat $(find ./report.log -mmin +10) | sed '/^Summary:/ s/$/ (cached)/' ---> Running in 3ca5eddaec86 ---> Removed intermediate container 3ca5eddaec86 ---> 56f0cb40a1f1 Step 47/47 : ENTRYPOINT [] ---> Running in 208af63ab4f5 ---> Removed intermediate container 208af63ab4f5 ---> 2c281b91cb05 [Warning] One or more build-args [GIT_COMMIT_SHA SPACK_CACHE] were not consumed Successfully built 2c281b91cb05 Successfully tagged us-central1-docker.pkg.dev/cp2k-org-project/cp2kci/img_cp2k-perf-cuda-volta:master Pushing new image... done. #################### Running Image cp2k-perf-cuda-volta #################### Uploading artifacts... done EndDate: 2026-10-10 02:23:52+00:00