StartDate: 2026-10-06 00:06:52+00:00 CpuId: 12x Intel Xeon W 2000 / D-2100 (Skylake / Cascade Lake) {Skylake}, 14nm GpuId: 1x Tesla V100-SXM2-16GB CommitSHA: 474e12a1c367a620a5548e7a1e77b007886d043e CommitTime: 2026-10-05 22:44:54 +0200 CommitAuthor: Dynamics of Condensed Matter CommitSubject: Move native Skala modules into src/xc (#6189) #################### Building Image cp2k-perf-cuda-volta #################### Dockerfile: /tools/docker/Dockerfile.test_performance_cuda_V100 Build-Path: / Build-Args: GIT_COMMIT_SHA=474e12a1c367a620a5548e7a1e77b007886d043e SPACK_CACHE=gs://cp2k-spack-cache Build-Cache: Yes Populating docker build cache... done. DEPRECATED: The legacy builder is deprecated and will be removed in a future release. BuildKit is currently disabled; enable it by removing the DOCKER_BUILDKIT=0 environment-variable. Sending build context to Docker daemon 501.3MB Step 1/47 : FROM docker.io/nvidia/cuda:12.9.1-devel-ubuntu24.04 12.9.1-devel-ubuntu24.04: Pulling from nvidia/cuda 32f112e3802c: Pulling fs layer 644e9b203583: Pulling fs layer 02559cd4bc8d: Pulling fs layer 2cd52cbb1ebe: Pulling fs layer 6e8af4fd0a07: Pulling fs layer 15a17189b2df: Pulling fs layer 02cb0e091e33: Pulling fs layer 9c3d619183d2: Pulling fs layer 7f7602a82106: Pulling fs layer 5a2aba542b08: Pulling fs layer 6cb9b761b877: Pulling fs layer 6e8af4fd0a07: Waiting 15a17189b2df: Waiting 02cb0e091e33: Waiting 9c3d619183d2: Waiting 7f7602a82106: Waiting 5a2aba542b08: Waiting 6cb9b761b877: Waiting 2cd52cbb1ebe: Waiting 644e9b203583: Verifying Checksum 644e9b203583: Download complete 32f112e3802c: Verifying Checksum 32f112e3802c: Download complete 2cd52cbb1ebe: Verifying Checksum 2cd52cbb1ebe: Download complete 6e8af4fd0a07: Verifying Checksum 6e8af4fd0a07: Download complete 02cb0e091e33: Verifying Checksum 02cb0e091e33: Download complete 9c3d619183d2: Verifying Checksum 9c3d619183d2: Download complete 7f7602a82106: Verifying Checksum 7f7602a82106: Download complete 02559cd4bc8d: Verifying Checksum 02559cd4bc8d: Download complete 6cb9b761b877: Verifying Checksum 6cb9b761b877: Download complete 32f112e3802c: Pull complete 644e9b203583: Pull complete 02559cd4bc8d: Pull complete 2cd52cbb1ebe: Pull complete 6e8af4fd0a07: Pull complete 15a17189b2df: Verifying Checksum 15a17189b2df: Download complete 5a2aba542b08: Verifying Checksum 5a2aba542b08: Download complete 15a17189b2df: Pull complete 02cb0e091e33: Pull complete 9c3d619183d2: Pull complete 7f7602a82106: Pull complete 5a2aba542b08: Pull complete 6cb9b761b877: Pull complete Digest: sha256:020bc241a628776338f4d4053fed4c38f6f7f3d7eb5919fecb8de313bb8ba47c Status: Downloaded newer image for nvidia/cuda:12.9.1-devel-ubuntu24.04 ---> eecafe98c3e1 Step 2/47 : ENV CUDA_PATH /usr/local/cuda ---> Using cache ---> 780681fb1fee Step 3/47 : ENV LD_LIBRARY_PATH /usr/local/cuda/lib64 ---> Using cache ---> ba98a15dc225 Step 4/47 : ENV CUDA_CACHE_DISABLE 1 ---> Using cache ---> 3932740340f7 Step 5/47 : RUN apt-get update -qq && apt-get install -qq --no-install-recommends gfortran && rm -rf /var/lib/apt/lists/* ---> Using cache ---> a06eb14abc29 Step 6/47 : WORKDIR /opt/cp2k-toolchain ---> Using cache ---> 082681bac850 Step 7/47 : COPY ./tools/toolchain/install_requirements*.sh ./ ---> Using cache ---> ae920e0abda3 Step 8/47 : RUN ./install_requirements.sh ubuntu ---> Using cache ---> 94839a704e2d Step 9/47 : RUN mkdir scripts ---> Using cache ---> 433a8b0a0499 Step 10/47 : COPY ./tools/toolchain/scripts/VERSION ./tools/toolchain/scripts/tool_kit.sh ./tools/toolchain/scripts/common_vars.sh ./tools/toolchain/scripts/signal_trap.sh ./scripts/ ---> Using cache ---> edd0ada5e677 Step 11/47 : COPY ./tools/toolchain/install_cp2k_toolchain.sh . ---> 3d8383e2bc04 Step 12/47 : RUN ./install_cp2k_toolchain.sh --with-mpich=install --mpi-mode=mpich --enable-cuda=yes --with-libgint=install --with-sirius=install --gpu-ver=V100 --dry-run ---> Running in 36f4c94335d5 No MPI installation detected. (Ignore this message if a fresh MPI installation is requested.) Toolchain script received the following options: --with-mpich=install --mpi-mode=mpich --enable-cuda=yes --with-libgint=install --with-sirius=install --gpu-ver=V100 --dry-run Parsing options and resolving conflicts... WARNING: (./install_cp2k_toolchain.sh, line 1121) Installing dependencies and CP2K requires CMake but CMake is not enabled, so a new copy of CMake will be installed first.  Toolchain configuration summary ------------------------------- System specifications: -j = 12 --target-cpu = native --gpu-ver = V100 --mpi-mode = mpich --math-mode = openblas Enabled features: --enable-tsan = no --enable-cuda = yes --enable-gauxc-cutlass = no --enable-hip = no --enable-opencl = no --enable-cray = no Packages to be installed: - cmake - mpich - openblas - fftw - eigen - libint - libxc - libxsmm - libxs - cosma - scalapack - elpa - dbcsr - spfft - spla - gsl - spglib - hdf5 - libvdwxc - sirius - libvori - tblite - pugixml - fmt - libgint Packages to be detected from system: - gcc Packages not used: - intel - amd - ninja - openmpi - intelmpi - mkl - acml - gauxc - libxstream - cusolvermp - plumed - libtorch - deepmd - ace - dftd4 - libsmeagol - trexio - libfci - greenx - gmp - mcl - skala_ftorch With --dry-run option, this script concludes with above report. The setup, toolchain env and conf files are written to /opt/cp2k-toolchain/install. ---> Removed intermediate container 36f4c94335d5 ---> a1ed72c96a3b Step 13/47 : COPY ./tools/toolchain/scripts/stage0/ ./scripts/stage0/ ---> 2ff1a9986c95 Step 14/47 : RUN ./scripts/stage0/install_stage0.sh && rm -rf ./build ---> Running in cc4c26c3e1c0 ==================== Finding GCC from system paths ==================== path to gcc is /usr/bin/gcc path to g++ is /usr/bin/g++ path to gfortran is /usr/bin/gfortran GCC compiler version 13.3.0 found Step gcc took 0.00 seconds. Step intel took 0.00 seconds. Step amd took 0.00 seconds. ==================== Installing CMake ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/cmake-4.3.0-linux-x86_64.tar.gz -O cmake-4.3.0-linux-x86_64.tar.gz cmake-4.3.0-linux-x86_64.tar.gz: OK Checksum of cmake-4.3.0-linux-x86_64.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/cmake-4.3.0 Step cmake took 8.00 seconds. Step ninja took 0.00 seconds. ---> Removed intermediate container cc4c26c3e1c0 ---> 1f42c06330a5 Step 15/47 : COPY ./tools/toolchain/scripts/stage1/ ./scripts/stage1/ ---> c2e7be511772 Step 16/47 : RUN ./scripts/stage1/install_stage1.sh && rm -rf ./build ---> Running in a6ab84e4c6b5 ==================== Installing MPICH ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/mpich-5.0.2.tar.gz -O mpich-5.0.2.tar.gz mpich-5.0.2.tar.gz: OK Checksum of mpich-5.0.2.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/mpich-5.0.2 for MPICH device ch4 Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/bin Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/lib Found directory /opt/cp2k-toolchain/install/mpich-5.0.2/include mpiexec is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpiexec mpicc is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpicc mpicxx is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpicxx mpifort is installed as /opt/cp2k-toolchain/install/mpich-5.0.2/bin/mpifort Step mpich took 581.00 seconds. ---> Removed intermediate container a6ab84e4c6b5 ---> ccc137aa314e Step 17/47 : COPY ./tools/toolchain/scripts/stage2/ ./scripts/stage2/ ---> eaab0102d054 Step 18/47 : RUN ./scripts/stage2/install_stage2.sh && rm -rf ./build ---> Running in a2c9334e5d10 ==================== Installing OpenBLAS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/OpenBLAS-0.3.34.tar.gz -O OpenBLAS-0.3.34.tar.gz OpenBLAS-0.3.34.tar.gz: OK Checksum of OpenBLAS-0.3.34.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/openblas-0.3.34 Installing OpenBLAS library for native target Step openblas took 304.00 seconds. Step gmp took 0.00 seconds. ---> Removed intermediate container a2c9334e5d10 ---> 33de1621d7b5 Step 19/47 : COPY ./tools/toolchain/scripts/stage3/ ./scripts/stage3/ ---> 24c3ef01ab76 Step 20/47 : RUN ./scripts/stage3/install_stage3.sh && rm -rf ./build ---> Running in f626b9117b18 ==================== Installing FFTW ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/fftw-3.3.11.tar.gz -O fftw-3.3.11.tar.gz fftw-3.3.11.tar.gz: OK Checksum of fftw-3.3.11.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/fftw-3.3.11 Step fftw took 169.00 seconds. ==================== Installing Eigen ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/eigen-5.0.1.tar.gz -O eigen-5.0.1.tar.gz eigen-5.0.1.tar.gz: OK Checksum of eigen-5.0.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/eigen-5.0.1 Step eigen took 3.00 seconds. ==================== Installing LIBINT ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libint-v2.13.1-cp2k-lmax-5.tar.xz -O libint-v2.13.1-cp2k-lmax-5.tar.xz libint-v2.13.1-cp2k-lmax-5.tar.xz: OK Checksum of libint-v2.13.1-cp2k-lmax-5.tar.xz Ok Installing from scratch into /opt/cp2k-toolchain/install/libint-v2.13.1-cp2k-lmax-5 Step libint took 537.00 seconds. ==================== Installing LIBXC ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxc-7.1.2.tar.bz2 -O libxc-7.1.2.tar.bz2 libxc-7.1.2.tar.bz2: OK Checksum of libxc-7.1.2.tar.bz2 Ok Installing from scratch into /opt/cp2k-toolchain/install/libxc-7.1.2 Installing CUDA-only libxc into /opt/cp2k-toolchain/install/libxc-7.1.2 Step libxc took 412.00 seconds. Step greenx took 0.00 seconds. ---> Removed intermediate container f626b9117b18 ---> 7dd8775a6eeb Step 21/47 : COPY ./tools/toolchain/scripts/stage4/ ./scripts/stage4/ ---> e6827566753a Step 22/47 : RUN ./scripts/stage4/install_stage4.sh && rm -rf ./build ---> Running in f3d3f68eb2b9 ==================== Installing Libxsmm ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxsmm-2.1.0.tar.gz -O libxsmm-2.1.0.tar.gz libxsmm-2.1.0.tar.gz: OK Checksum of libxsmm-2.1.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libxsmm-2.1.0 Step libxsmm took 22.00 seconds. ==================== Installing LIBXS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libxs-1.0.0.tar.gz -O libxs-1.0.0.tar.gz libxs-1.0.0.tar.gz: OK Checksum of libxs-1.0.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libxs-1.0.0 Step libxs took 8.00 seconds. Step libxstream took 0.00 seconds. ==================== Installing libGint ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libGint-v1.tar.gz -O libGint-v1.tar.gz libGint-v1.tar.gz: OK Checksum of libGint-v1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libGint-v1 Step libGint took 123.00 seconds. ==================== Installing ScaLAPACK ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/scalapack-2.2.3.tar.gz -O scalapack-2.2.3.tar.gz scalapack-2.2.3.tar.gz: OK Checksum of scalapack-2.2.3.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/scalapack-2.2.3 Step scalapack took 37.00 seconds. Step cusolvermp took 0.00 seconds. ==================== Installing COSMA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/COSMA-v2.8.4.tar.gz -O COSMA-v2.8.4.tar.gz COSMA-v2.8.4.tar.gz: OK Checksum of COSMA-v2.8.4.tar.gz Ok wget --tries=5 --quiet https://www.cp2k.org/static/downloads/COSTA-v2.3.2.tar.gz -O COSTA-v2.3.2.tar.gz COSTA-v2.3.2.tar.gz: OK Checksum of COSTA-v2.3.2.tar.gz Ok wget --tries=5 --quiet https://www.cp2k.org/static/downloads/Tiled-MM-v2.3.2.tar.gz -O Tiled-MM-v2.3.2.tar.gz Tiled-MM-v2.3.2.tar.gz: OK Checksum of Tiled-MM-v2.3.2.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/COSMA-2.8.4 Step cosma took 67.00 seconds. ---> Removed intermediate container f3d3f68eb2b9 ---> 39a8664e960b Step 23/47 : COPY ./tools/toolchain/scripts/stage5/ ./scripts/stage5/ ---> c17c7865b2e0 Step 24/47 : RUN ./scripts/stage5/install_stage5.sh && rm -rf ./build ---> Running in e4be38138cc2 ==================== Installing ELPA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/elpa-2026.02.002.tar.gz -O elpa-2026.02.002.tar.gz elpa-2026.02.002.tar.gz: OK Checksum of elpa-2026.02.002.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002 Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002/cpu Installing from scratch into /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia Step elpa took 319.00 seconds. Step skala took 0.00 seconds. ---> Removed intermediate container e4be38138cc2 ---> adb8e3dd29bd Step 25/47 : COPY ./tools/toolchain/scripts/stage6/ ./scripts/stage6/ ---> 8bde518f265f Step 26/47 : RUN ./scripts/stage6/install_stage6.sh && rm -rf ./build ---> Running in dbcd7b5e10d7 ==================== Installing GSL ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/gsl-2.8.tar.gz -O gsl-2.8.tar.gz gsl-2.8.tar.gz: OK Checksum of gsl-2.8.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/gsl-2.8 Step gsl took 75.00 seconds. Step plumed took 0.00 seconds. Step libtorch took 0.00 seconds. Step ftorch took 0.00 seconds. Step gauxc took 0.00 seconds. Step deepmd took 0.00 seconds. Step ace took 0.00 seconds. ---> Removed intermediate container dbcd7b5e10d7 ---> a49d8a6ae877 Step 27/47 : COPY ./tools/toolchain/scripts/stage7/ ./scripts/stage7/ ---> 2b145b350e3f Step 28/47 : RUN ./scripts/stage7/install_stage7.sh && rm -rf ./build ---> Running in 4ed0c84ec203 ==================== Installing HDF5 ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/hdf5-2.2.0.tar.gz -O hdf5-2.2.0.tar.gz hdf5-2.2.0.tar.gz: OK Checksum of hdf5-2.2.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/hdf5-2.2.0 Step hdf5 took 130.00 seconds. ==================== Installing libvdwxc ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libvdwxc-0.5.0.tar.gz -O libvdwxc-0.5.0.tar.gz libvdwxc-0.5.0.tar.gz: OK Checksum of libvdwxc-0.5.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libvdwxc-0.5.0 Step libvdwxc took 15.00 seconds. ==================== Installing Spglib ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/spglib-2.7.0.tar.gz -O spglib-2.7.0.tar.gz spglib-2.7.0.tar.gz: OK Checksum of spglib-2.7.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/spglib-2.7.0 Step spglib took 5.00 seconds. ==================== Installing libvori ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/libvori-220621.tar.gz -O libvori-220621.tar.gz libvori-220621.tar.gz: OK Checksum of libvori-220621.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/libvori-220621 Step libvori took 13.00 seconds. Step libsmeagol took 0.00 seconds. Step libfci took 0.00 seconds. ==================== Installing fmt ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/fmt-12.1.0.zip -O fmt-12.1.0.zip fmt-12.1.0.zip: OK Checksum of fmt-12.1.0.zip Ok Installing from scratch into /opt/cp2k-toolchain/install/fmt-12.1.0 Step fmt took 8.00 seconds. ---> Removed intermediate container 4ed0c84ec203 ---> 84e8cf587987 Step 29/47 : COPY ./tools/toolchain/scripts/stage8/ ./scripts/stage8/ ---> 449ed7642bb9 Step 30/47 : RUN ./scripts/stage8/install_stage8.sh && rm -rf ./build ---> Running in 2818100c4c18 Step dftd4 took 0.00 seconds. ==================== Installing tblite ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/tblite-0.7.0.tar.xz -O tblite-0.7.0.tar.xz tblite-0.7.0.tar.xz: OK Checksum of tblite-0.7.0.tar.xz Ok Step tblite took 43.00 seconds. ==================== Installing pugixml ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/pugixml-1.15.tar.gz -O pugixml-1.15.tar.gz pugixml-1.15.tar.gz: OK Checksum of pugixml-1.15.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/pugixml-1.15 Step pugixml took 10.00 seconds. ==================== Installing SpFFT ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SpFFT-1.1.1.tar.gz -O SpFFT-1.1.1.tar.gz SpFFT-1.1.1.tar.gz: OK Checksum of SpFFT-1.1.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/SpFFT-1.1.1 Step spfft took 21.00 seconds. ==================== Installing SpLA ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SpLA-1.6.1.tar.gz -O SpLA-1.6.1.tar.gz SpLA-1.6.1.tar.gz: OK Checksum of SpLA-1.6.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/SpLA-1.6.1 Step spla took 24.00 seconds. ==================== Installing SIRIUS ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/SIRIUS-7.11.1.tar.gz -O SIRIUS-7.11.1.tar.gz SIRIUS-7.11.1.tar.gz: OK Checksum of SIRIUS-7.11.1.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/sirius-7.11.1 Installing from scratch into /opt/cp2k-toolchain/install/sirius-7.11.1/cuda Step sirius took 443.00 seconds. Step trexio took 0.00 seconds. Step MCL took 0.00 seconds. ---> Removed intermediate container 2818100c4c18 ---> cce50f714e34 Step 31/47 : COPY ./tools/toolchain/scripts/stage9/ ./scripts/stage9/ ---> 12c8dd6d6080 Step 32/47 : RUN ./scripts/stage9/install_stage9.sh && rm -rf ./build ---> Running in dcf979ead41d ==================== Installing DBCSR ==================== wget --tries=5 --quiet https://www.cp2k.org/static/downloads/dbcsr-2.10.0.tar.gz -O dbcsr-2.10.0.tar.gz dbcsr-2.10.0.tar.gz: OK Checksum of dbcsr-2.10.0.tar.gz Ok Installing from scratch into /opt/cp2k-toolchain/install/dbcsr-2.10.0 Installing from scratch into /opt/cp2k-toolchain/install/dbcsr-2.10.0-cuda Step DBCSR took 127.00 seconds. ---> Removed intermediate container dcf979ead41d ---> dd391319bf8b Step 33/47 : WORKDIR /opt/cp2k ---> Running in e56fae1df62b ---> Removed intermediate container e56fae1df62b ---> 2b15d84dfaa2 Step 34/47 : COPY ./src ./src ---> ad40982f4bdb Step 35/47 : COPY ./data ./data ---> 594742dd8a92 Step 36/47 : COPY ./tools/build_utils ./tools/build_utils ---> 10ef9689f62f Step 37/47 : COPY ./cmake ./cmake ---> f4cadbcdde72 Step 38/47 : COPY ./CMakeLists.txt . ---> 6477c0ff581d Step 39/47 : COPY ./CMakePresets.json . ---> f9dbabe2599e Step 40/47 : COPY ./tools/docker/scripts/build_cp2k.sh ./tools/docker/scripts/cmake_cp2k.sh ./ ---> 5d0919e1f2ae Step 41/47 : RUN ./build_cp2k.sh toolchain_cuda_V100 psmp ---> Running in a58a8d0217f2 ==================== Building CP2K ==================== -- The Fortran compiler identification is GNU 13.3.0 -- The C compiler identification is GNU 13.3.0 -- The CXX compiler identification is GNU 13.3.0 -- Detecting Fortran compiler ABI info -- Detecting Fortran compiler ABI info - done -- Check for working Fortran compiler: /usr/bin/gfortran - skipped -- Detecting C compiler ABI info -- Detecting C compiler ABI info - done -- Check for working C compiler: /usr/bin/gcc - skipped -- Detecting C compile features -- Detecting C compile features - done -- Detecting CXX compiler ABI info -- Detecting CXX compiler ABI info - done -- Check for working CXX compiler: /usr/bin/g++ - skipped -- Detecting CXX compile features -- Detecting CXX compile features - done -- Found PkgConfig: /usr/bin/pkg-config (found version "1.8.1") -- Found Python: /usr/bin/python3.12 (found version "3.12.3") found components: Interpreter -- Found MPI_C: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpi.so (found version "5.0") -- Found MPI_CXX: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpicxx.so (found version "5.0") -- Found MPI_Fortran: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpifort.so (found version "5.0") -- Found MPI: TRUE (found version "5.0") found components: C CXX Fortran -- Could NOT find MKL (missing: CP2K_MKL_INCLUDE_DIRS _mkl_interface_library _mkl_thread_library _mkl_core_library _mkl_scalapack_library _mkl_blacs_library) -- Checking for module 'openblas' -- Found openblas, version 0.3.34 -- Found OpenBLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/include -- Found Blas: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found Lapack: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Checking for module 'scalapack' -- Package 'mpi', required by 'scalapack', not found Package 'lapack', required by 'scalapack', not found Package 'blas', required by 'scalapack', not found -- Found SCALAPACK: /opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a -- Found Threads: TRUE -- Using LIBXS + LIBXSMM for Small Matrix Multiplication -- CP2K_WITH_GPU is deprecated in favor of CMAKE_HIP_ARCHITECTURES or CMAKE_CUDA_ARCHITECTURES ------------------------------------------------------------ - DBCSR - ------------------------------------------------------------ -- Found MPI: TRUE (found version "5.0") -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") -- The CUDA compiler identification is NVIDIA 12.9.86 with host compiler GNU 13.3.0 -- Detecting CUDA compiler ABI info -- Detecting CUDA compiler ABI info - done -- Check for working CUDA compiler: /usr/local/cuda/bin/nvcc - skipped -- Detecting CUDA compile features -- Detecting CUDA compile features - done -- Found CUDAToolkit: /usr/local/cuda/targets/x86_64-linux/include (found version "12.9.86") ----------------------------------------------------------- - CUDA - ----------------------------------------------------------- -- GPU architecture number: 70 -- GPU profiling enabled: OFF -- CUDA compiler and libraries found ------------------------------------------------------------ - OPENMP - ------------------------------------------------------------ -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") found components: Fortran C CXX ------------------------------------------------------------ - Other dependencies - ------------------------------------------------------------ -- Checking for one of the modules 'elpa_openmp' -- Found Elpa: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so;cudart;cublasLt;cublas;/opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a;:libopenblas.a -- Found CUDA Libxc (7.1.2) -- Found HDF5: hdf5-shared;hdf5_fortran-shared (found version "2.2.0") found components: C Fortran -- Found MPI: TRUE (found version "5.0") found components: CXX -- Found OPENBLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found Blas: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Checking for one of the modules 'fftw3' -- Checking for one of the modules 'fftw3f' -- Checking for one of the modules 'fftw3l' -- Checking for one of the modules 'fftw3q' -- Found Fftw: /opt/cp2k-toolchain/install/fftw-3.3.11/include -- Boost detected. satisfied by headers bundled with Libint2 distribution -- Found LibGint: /opt/cp2k-toolchain/install/libGint-v1/lib/libcp2kGint.a -- Component omp of Spglib: NOT FOUND -- Component fortran of Spglib: FOUND (LIB_TYPE: static) -- Found package: Spglib -- Looking for Fortran sgemm -- Looking for Fortran sgemm - found -- multicharge: Find installed package -- toml-f: Find installed package -- s-dftd3: Find installed package -- Found GSL: /opt/cp2k-toolchain/install/gsl-2.8/include (found version "2.8") -- Checking for one of the modules 'libxc>=3.0.0' -- Found LibXC: /opt/cp2k-toolchain/install/libxc-7.1.2/lib/libxc.so (Required is at least version "3.0.0") -- Found LibSPG: /opt/cp2k-toolchain/install/spglib-2.7.0/lib/libsymspg.a -- Found HDF5: hdf5-shared (found version "2.2.0") found components: C -- Found FFTW: /opt/cp2k-toolchain/install/fftw-3.3.11/include -- Looking for Fortran sgemm -- Looking for Fortran sgemm - not found -- Found BLAS: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP_CUDA: -fopenmp (found version "4.5") -- Found OpenMP_Fortran: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") -- Checking for one of the modules 's-dftd3' -- Checking for one of the modules 'mctc-lib' -- Found DFTD3: /opt/cp2k-toolchain/install/tblite-0.7.0/lib/libs-dftd3.a -- Checking for one of the modules 'dftd4' -- Checking for one of the modules 'multicharge' -- Found DFTD4: /opt/cp2k-toolchain/install/tblite-0.7.0/lib/libdftd4.a -- Looking for Fortran cheev -- Looking for Fortran cheev - found -- Found LAPACK: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so;-lm;-ldl -- Checking for one of the modules 'scalapack' -- Checking for one of the modules 'elpa;elpa_openmp;elpa-openmp-2019.05.001;elpa_openmp-2019.11.001;elpa_openmp-2020.05.001;elpa-2019.05.001;elpa-2019.11.001;elpa-2020.05.001' -- Found Elpa: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so -- Checking for module 'libvdwxc>=0.5.0' -- Found libvdwxc, version 0.5.0 -- Checking for module 'fftw3' -- Found fftw3, version 3.3.11 -- Found LibVDWXC: vdwxc;fftw3 (Required is at least version "0.5.0") -- Setting build type to 'Release' as none was specified. -- Performing Test f2008-norm2 -- Performing Test f2008-norm2 - Success -- Performing Test f2008-block_construct -- Performing Test f2008-block_construct - Success -- Performing Test f2008-contiguous -- Performing Test f2008-contiguous - Success -- Performing Test f2008-findloc -- Performing Test f2008-findloc - Success -- Performing Test f95-reshape-order-allocatable -- Performing Test f95-reshape-order-allocatable - Success -- FYPP preprocessor found. -- Adding libxs_jit.F from dependency libxs for compilation -------------------------------------------------------------------- - - - Summary of enabled dependencies - - - -------------------------------------------------------------------- - BLAS - Vendor: OpenBLAS - Include directories: /opt/cp2k-toolchain/install/openblas-0.3.34/include - Libraries: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so - LAPACK - Include directories: /opt/cp2k-toolchain/install/openblas-0.3.34/include - Libraries: /opt/cp2k-toolchain/install/openblas-0.3.34/lib/libopenblas.so - FFTW3 - Include directories: /opt/cp2k-toolchain/install/fftw-3.3.11/include - Libraries: /opt/cp2k-toolchain/install/fftw-3.3.11/lib/libfftw3.a - MPI - Include directories: /opt/cp2k-toolchain/install/mpich-5.0.2/include - Libraries: /opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpicxx.so;/opt/cp2k-toolchain/install/mpich-5.0.2/lib/libmpi.so - MPI_F08: Enabled - ScaLAPACK - Vendor: auto - Include directories: - Libraries: /opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a - Hardware acceleration - Backend: CUDA - GPU architectures: 70 - GPU profiling enabled: OFF - GPU-accelerated modules - ELPA: ON - GRID: ON - DBM: ON - PW: ON - LIBXC: ON - LibXC - Include directories: LIBXC_CPU_INCLUDE_DIRS-NOTFOUND - Libraries: Libxc::xc - Spglib - Include directories: /opt/cp2k-toolchain/install/spglib-2.7.0/include;$ - HDF5 - Include directories: /opt/cp2k-toolchain/install/hdf5-2.2.0/include - Libraries: hdf5-shared - LIBXS - Include directories: - Libraries: - SpLA - Include directories: /opt/cp2k-toolchain/install/SpLA-1.6.1-cuda/include;/opt/cp2k-toolchain/install/SpLA-1.6.1-cuda/include/spla - Libraries: $;$;$;$;MPI::MPI_CXX;MPI::MPI_C;MPI::MPI_Fortran - SpLA GEMM offloading - DFTD4 - Enabled via TBLITE - Include directories: /opt/cp2k-toolchain/install/tblite-0.7.0/include;/opt/cp2k-toolchain/install/tblite-0.7.0/include/dftd4/GNU-13.3.0 - Libraries: - TBLITE - Include directories: - Libraries: - SIRIUS - Include directories: - Libraries: - COSMA - Include directories: /opt/cp2k-toolchain/install/COSMA-2.8.4-cuda/include - Libraries: MPI::MPI_CXX;costa::costa;$;$;$<$:cosma::BLAS::blas>;$;$<$:Tiled-MM::Tiled-MM>;$<$:Tiled-MM::Tiled-MM>;$<$:semiprof::semiprof>;$<$:cosma::scalapack::scalapack> - Libint2 - Include directories: - Libraries: - LibGint - include directories: /opt/cp2k-toolchain/install/libGint-v1/include - libraries: /opt/cp2k-toolchain/install/libGint-v1/lib/libcp2kGint.a - ELPA - Include directories: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/include/elpa_openmp-2026.02.002 - Libraries: /opt/cp2k-toolchain/install/elpa-2026.02.002/nvidia/lib/libelpa_openmp.so;cudart;cublasLt;cublas;/opt/cp2k-toolchain/install/scalapack-2.2.3/lib/libscalapack.a;:libopenblas.a -------------------------------------------------------------------- - - - Dependencies not included in this build - - - -------------------------------------------------------------------- - DeePMD - PEXSI - ACE (libpace) - LibSMEAGOL - MiMiC - DLA-Future - PLUMED - LibFCI - GauXC - Skala/FTorch - Libvori - LibTorch - TREXIO - OpenPMD - GreenX After building and installing CP2K, run the regtests with: /opt/cp2k/tests/do_regtest.py /opt/cp2k/bin psmp -- Configuring done (13.9s) -- Generating done (0.9s) -- Build files have been written to: /opt/cp2k/build Compiling CP2K ... done ---> Removed intermediate container a58a8d0217f2 ---> 7e788102624c Step 42/47 : COPY ./benchmarks ./benchmarks ---> ac763fd56b7a Step 43/47 : COPY ./tools/regtesting ./tools/regtesting ---> c2b40c2da2d2 Step 44/47 : COPY ./tools/docker/scripts/test_performance.sh ./tools/docker/scripts/plot_performance.py ./ ---> 00f289849bc3 Step 45/47 : RUN ./test_performance.sh "toolchain_cuda_V100" 2>&1 | tee report.log ---> Running in e30946923835 ============== CP2K Binary Flags ============= cp2kflags: omp libint fftw3 libxc libxc_gpu elpa parallel scalapack mpi_f08 cosma libxs libxsmm dbcsr_acc spglib openblas libdftd4 s_dftd3 mctc-lib tblite sirius offload_cuda spla_gemm_offloading libvdwxc hdf5 libGint ========== Checking Benchmark Inputs ========= Found 89 input files and 0 errors. ========== Running Performance Test ========== Plot: name="total_timings_6cpu_1gpu", title="Total Timings with 6 CPU Cores and 1 GPU", ylabel="time [s]" Running H2O-64.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/H2O-64_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.030 0.031 106.432 106.432 qs_mol_dyn_low 1 2.0 0.005 0.005 105.987 105.990 qs_forces 11 3.9 0.002 0.002 105.935 105.935 qs_energies 11 4.9 0.002 0.002 94.587 94.589 scf_env_do_scf 11 5.9 0.001 0.001 79.046 79.046 velocity_verlet 10 3.0 0.002 0.002 66.266 66.285 scf_env_do_scf_inner_loop 108 6.5 0.007 0.010 64.804 64.804 rebuild_ks_matrix 119 8.3 0.001 0.001 27.697 27.698 qs_ks_build_kohn_sham_matrix 119 9.3 0.027 0.027 27.696 27.697 dbcsr_multiply_generic 2319 12.5 0.159 0.159 26.571 26.585 qs_ks_update_qs_env 119 7.6 0.001 0.001 25.720 25.723 qs_rho_update_rho_low 119 7.7 0.001 0.001 22.166 22.185 calculate_rho_elec 119 8.7 0.897 0.905 22.165 22.184 qs_scf_new_mos 108 7.5 0.001 0.001 21.997 22.013 qs_scf_loop_do_ot 108 8.5 0.001 0.001 21.997 22.012 ot_scf_mini 108 9.5 0.003 0.003 19.952 19.952 fft_wrap_pw1pw2 1201 11.6 0.025 0.025 17.029 17.049 fft_wrap_pw1pw2_140 487 12.2 0.003 0.003 14.638 14.655 sum_up_and_integrate 119 10.3 0.005 0.005 14.506 14.540 integrate_v_rspace 119 11.3 0.365 0.366 14.407 14.441 init_scf_loop 11 6.9 0.001 0.001 14.179 14.179 multiply_cannon 2319 13.5 0.365 0.370 13.333 13.347 multiply_cannon_loop 2319 14.5 0.291 0.292 12.165 12.202 ot_mini 108 10.5 0.001 0.001 11.571 11.573 make_m2s 4638 13.5 0.048 0.048 11.467 11.480 make_images 4638 14.5 1.179 1.194 11.281 11.294 density_rs2pw 119 9.7 0.008 0.008 11.134 11.293 prepare_preconditioner 11 7.9 0.000 0.000 10.819 10.820 make_preconditioner 11 8.9 0.000 0.000 10.819 10.820 grid_collocate_task_list 119 9.7 10.105 10.241 10.105 10.241 make_full_inverse_cholesky 11 9.9 0.002 0.002 9.704 9.982 pw_gpu_r3dc1d_3d_ps 606 13.1 2.408 2.413 8.711 8.724 pw_gpu_c1dr3d_3d_ps 595 14.2 2.286 2.301 8.286 8.319 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 8.087 8.213 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 7.885 7.885 grid_integrate_task_list 119 12.3 7.454 7.487 7.454 7.487 qs_ot_get_derivative 108 11.5 0.002 0.002 7.123 7.126 init_scf_run 11 5.9 0.000 0.000 6.992 6.993 scf_env_initial_rho_setup 11 6.9 0.001 0.001 6.992 6.992 make_images_data 4638 15.5 0.062 0.062 6.692 6.716 potential_pw2rs 119 12.3 0.037 0.038 6.587 6.588 multiply_cannon_multrec 4638 15.5 2.076 2.078 6.520 6.528 hybrid_alltoall_any 4638 16.5 4.983 4.984 6.448 6.473 copy_dbcsr_to_fm 153 11.3 0.147 0.148 5.801 5.804 transfer_dbcsr_to_fm 11 10.9 0.033 0.036 5.392 5.393 dbcsr_to_fm_plan_create 11 12.9 4.537 4.605 5.059 5.062 ot_diis_step 108 11.5 0.007 0.007 4.424 4.424 mp_alltoall_z22v 1201 15.6 4.405 4.405 4.405 4.405 build_core_ppl_forces 11 5.9 4.055 4.146 4.055 4.146 build_core_hamiltonian_matrix 11 6.9 0.001 0.001 4.038 4.086 wfi_extrapolate 11 7.9 0.001 0.001 4.060 4.060 dbcsr_mm_accdrv_process 9638 16.2 0.807 0.951 4.029 4.030 mp_waitall_1 65419 16.9 3.821 3.872 3.821 3.872 apply_preconditioner_dbcsr 119 12.6 0.000 0.000 3.822 3.824 apply_single 119 13.6 0.001 0.001 3.822 3.824 qs_ot_get_p 119 10.4 0.001 0.002 3.682 3.688 calculate_dm_sparse 119 9.5 0.001 0.001 3.522 3.539 qs_env_update_s_mstruct 11 6.9 0.000 0.000 3.475 3.488 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 3.124 3.125 multiply_cannon_sync_h2d 4638 15.5 3.103 3.105 3.103 3.105 qs_ot_get_derivative_taylor 59 13.0 0.003 0.003 2.929 2.929 transfer_rs2pw 487 10.6 0.009 0.009 2.719 2.922 cp_dbcsr_sm_fm_multiply 37 9.5 0.001 0.001 2.772 2.772 jit_kernel_multiply 12 15.9 2.607 2.749 2.607 2.749 pw_poisson_solve 119 10.3 0.003 0.003 2.726 2.730 yz_to_x 606 14.1 0.473 0.476 2.704 2.708 x_to_yz 595 15.2 0.512 0.515 2.686 2.688 qs_create_task_list 11 7.9 0.000 0.000 2.600 2.665 generate_qs_task_list 11 8.9 1.191 1.204 2.600 2.665 transfer_rs2pw_140 130 11.5 1.571 1.593 2.261 2.473 calculate_first_density_matrix 1 7.0 0.000 0.000 2.471 2.471 cp_fm_cholesky_invert 11 10.9 2.390 2.390 2.390 2.390 qs_ot_p2m_diag 50 11.0 0.089 0.090 2.381 2.382 cp_dbcsr_sm_fm_multiply_core 37 10.5 0.000 0.000 2.310 2.311 pw_gpu_fg 606 14.1 2.183 2.215 2.183 2.215 dbcsr_complete_redistribute 318 12.2 0.751 0.760 1.893 2.168 dbcsr_special_finalize 6957 15.5 0.043 0.043 2.151 2.160 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="H2O-64", label="H2O-64", y=106.432, yerr=0.0 Plot: name="H2O-64_timings_6cpu_1gpu", title="Timings of H2O-64 with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="rest", label="rest", y=74.94800000000001, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=10.105, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=7.454, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=4.983, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=4.537, yerr=0.0 PlotPoint: plot="H2O-64_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=4.405, yerr=0.0 Running H2O-64_nonortho.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/H2O-64_nonortho_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.030 0.032 97.277 97.277 qs_mol_dyn_low 1 2.0 0.004 0.004 96.814 96.817 qs_forces 11 3.9 0.002 0.002 96.765 96.765 qs_energies 11 4.9 0.001 0.001 85.392 85.393 scf_env_do_scf 11 5.9 0.001 0.001 69.300 69.300 velocity_verlet 10 3.0 0.002 0.002 61.960 61.978 scf_env_do_scf_inner_loop 96 6.5 0.006 0.008 55.289 55.289 rebuild_ks_matrix 107 8.3 0.001 0.001 25.276 25.276 qs_ks_build_kohn_sham_matrix 107 9.3 0.023 0.024 25.275 25.276 dbcsr_multiply_generic 1999 12.5 0.135 0.135 23.682 23.709 qs_ks_update_qs_env 107 7.6 0.001 0.001 23.130 23.130 qs_scf_new_mos 96 7.5 0.001 0.001 19.219 19.229 qs_scf_loop_do_ot 96 8.5 0.001 0.001 19.219 19.228 qs_rho_update_rho_low 107 7.7 0.001 0.001 17.765 17.774 calculate_rho_elec 107 8.7 0.801 0.804 17.765 17.773 ot_scf_mini 96 9.5 0.003 0.003 17.419 17.422 fft_wrap_pw1pw2 1081 11.6 0.021 0.021 15.371 15.410 init_scf_loop 11 6.9 0.001 0.001 13.948 13.948 sum_up_and_integrate 107 10.3 0.004 0.004 13.559 13.571 integrate_v_rspace 107 11.3 0.323 0.323 13.471 13.483 fft_wrap_pw1pw2_140 439 12.2 0.003 0.003 13.214 13.240 multiply_cannon 1999 13.5 0.307 0.311 11.947 11.964 multiply_cannon_loop 1999 14.5 0.250 0.254 10.989 10.996 prepare_preconditioner 11 7.9 0.000 0.000 10.614 10.620 make_preconditioner 11 8.9 0.000 0.000 10.614 10.620 make_m2s 3998 13.5 0.042 0.042 10.163 10.167 ot_mini 96 10.5 0.001 0.001 10.153 10.153 density_rs2pw 107 9.7 0.007 0.008 9.967 10.064 make_images 3998 14.5 1.061 1.073 10.000 10.004 make_full_inverse_cholesky 11 9.9 0.002 0.002 9.457 9.734 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 8.803 8.803 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 8.046 8.194 pw_gpu_r3dc1d_3d_ps 546 13.1 2.155 2.168 7.895 7.897 pw_gpu_c1dr3d_3d_ps 535 14.2 2.038 2.057 7.449 7.486 grid_integrate_task_list 107 12.3 7.213 7.226 7.213 7.226 grid_collocate_task_list 107 9.7 6.967 7.037 6.967 7.037 init_scf_run 11 5.9 0.000 0.000 6.640 6.640 scf_env_initial_rho_setup 11 6.9 0.000 0.001 6.639 6.639 qs_ot_get_derivative 96 11.5 0.002 0.002 6.286 6.286 multiply_cannon_multrec 3998 15.5 1.856 1.865 6.022 6.038 make_images_data 3998 15.5 0.052 0.053 5.935 5.943 potential_pw2rs 107 12.3 0.033 0.034 5.935 5.935 hybrid_alltoall_any 3998 16.5 4.427 4.429 5.723 5.731 copy_dbcsr_to_fm 147 11.2 0.143 0.144 5.529 5.529 transfer_dbcsr_to_fm 11 10.9 0.035 0.042 5.133 5.139 dbcsr_to_fm_plan_create 11 12.9 4.436 4.470 4.774 4.786 qs_env_update_s_mstruct 11 6.9 0.000 0.000 4.342 4.457 build_core_ppl_forces 11 5.9 4.009 4.119 4.009 4.119 build_core_hamiltonian_matrix 11 6.9 0.001 0.001 4.045 4.075 mp_alltoall_z22v 1081 15.6 3.905 3.940 3.905 3.940 ot_diis_step 96 11.5 0.005 0.005 3.844 3.844 dbcsr_mm_accdrv_process 8494 16.1 0.764 0.918 3.810 3.826 wfi_extrapolate 11 7.9 0.001 0.001 3.792 3.792 qs_create_task_list 11 7.9 0.000 0.000 3.484 3.568 generate_qs_task_list 11 8.9 1.493 1.509 3.483 3.568 apply_preconditioner_dbcsr 107 12.6 0.000 0.000 3.401 3.403 apply_single 107 13.6 0.001 0.001 3.401 3.403 mp_waitall_1 56411 16.9 3.317 3.326 3.317 3.326 calculate_dm_sparse 107 9.5 0.001 0.001 3.233 3.241 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 3.172 3.172 qs_ot_get_p 107 10.4 0.001 0.001 3.155 3.157 multiply_cannon_sync_h2d 3998 15.5 2.786 2.819 2.786 2.819 cp_dbcsr_sm_fm_multiply 37 9.5 0.001 0.001 2.697 2.698 jit_kernel_multiply 12 15.8 2.502 2.671 2.502 2.671 transfer_rs2pw 439 10.6 0.008 0.008 2.411 2.520 qs_ot_get_derivative_taylor 53 13.0 0.003 0.003 2.514 2.514 pw_poisson_solve 107 10.3 0.003 0.003 2.428 2.430 yz_to_x 546 14.1 0.424 0.427 2.413 2.428 calculate_first_density_matrix 1 7.0 0.000 0.000 2.413 2.414 x_to_yz 535 15.2 0.454 0.455 2.370 2.385 cp_fm_cholesky_invert 11 10.9 2.371 2.371 2.371 2.371 cp_dbcsr_sm_fm_multiply_core 37 10.5 0.000 0.000 2.241 2.242 dbcsr_complete_redistribute 306 12.1 0.749 0.790 1.925 2.196 transfer_rs2pw_140 118 11.5 1.429 1.442 2.004 2.119 build_core_ppl 11 7.9 2.054 2.078 2.054 2.078 qs_ot_p2m_diag 44 11.0 0.078 0.079 2.048 2.050 pw_gpu_fg 546 14.1 2.026 2.029 2.026 2.029 build_overlap_matrix_low 22 6.9 1.913 1.921 2.005 2.013 copy_fm_to_dbcsr 170 11.1 0.002 0.002 1.720 1.998 build_kinetic_matrix_low 22 6.9 1.886 1.891 1.985 1.990 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="H2O-64_nonortho", label="H2O-64_nonortho", y=97.277, yerr=0.0 Plot: name="H2O-64_nonortho_timings_6cpu_1gpu", title="Timings of H2O-64_nonortho with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="rest", label="rest", y=70.225, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=7.213, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=6.967, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=4.436, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=4.427, yerr=0.0 PlotPoint: plot="H2O-64_nonortho_timings_6cpu_1gpu", name="build_core_ppl_forces", label="build_core_ppl_forces", y=4.009, yerr=0.0 Running w64PBE.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/w64PBE_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.043 0.045 242.456 242.456 qs_mol_dyn_low 1 2.0 0.005 0.005 241.718 241.720 qs_forces 11 3.9 0.002 0.002 241.660 241.660 qs_energies 11 4.9 0.001 0.002 209.782 209.783 velocity_verlet 10 3.0 0.002 0.002 190.734 190.755 scf_env_do_scf 11 5.9 0.001 0.002 188.451 188.452 scf_env_do_scf_inner_loop 106 6.8 0.006 0.009 162.793 162.793 rebuild_ks_matrix 117 8.5 0.001 0.001 120.933 120.938 qs_ks_build_kohn_sham_matrix 117 9.5 0.026 0.027 120.932 120.937 qs_ks_update_qs_env 120 7.8 0.001 0.001 107.440 107.446 fft_wrap_pw1pw2 2000 12.9 0.049 0.050 70.691 70.699 fft_wrap_pw1pw2_200 1298 14.3 0.009 0.009 66.989 67.034 qs_rho_update_rho_low 117 7.9 0.001 0.001 61.857 61.865 calculate_rho_elec 117 8.9 1.255 1.256 61.857 61.864 qs_vxc_create 117 10.5 0.002 0.003 61.594 61.595 xc_vxc_pw_create 117 11.5 1.503 1.506 61.591 61.593 sum_up_and_integrate 117 10.5 0.005 0.005 44.806 44.907 integrate_v_rspace 117 11.5 0.219 0.220 44.615 44.716 grid_collocate_task_list 117 9.9 41.901 42.050 41.901 42.050 xc_pw_derive 702 13.5 0.010 0.010 39.280 39.320 pw_gpu_c1dr3d_3d_ps 1053 15.2 10.747 10.788 37.792 37.813 xc_rho_set_and_dset_create 117 12.5 0.937 0.941 33.523 33.544 grid_integrate_task_list 117 12.5 33.193 33.295 33.193 33.295 pw_gpu_r3dc1d_3d_ps 947 14.5 9.736 9.759 32.838 32.868 xc_pw_divergence 117 12.5 0.006 0.006 26.148 26.150 init_scf_loop 14 6.8 0.001 0.001 25.608 25.608 mp_alltoall_z22v 2000 16.9 19.057 19.187 19.057 19.187 density_rs2pw 117 9.9 0.010 0.010 18.663 18.827 dbcsr_multiply_generic 2077 12.5 0.150 0.151 18.565 18.662 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 17.428 17.535 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 14.297 14.298 qs_scf_new_mos 106 7.8 0.001 0.001 13.904 13.910 qs_scf_loop_do_ot 106 8.8 0.001 0.001 13.903 13.909 x_to_yz 1053 16.2 2.508 2.509 12.597 12.690 ot_scf_mini 106 9.8 0.003 0.003 12.464 12.469 xc_functional_eval 117 13.5 0.002 0.002 12.213 12.221 pbe_lda_eval 117 14.5 12.212 12.220 12.212 12.220 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 11.836 11.836 potential_pw2rs 117 12.5 0.060 0.060 11.204 11.207 yz_to_x 947 15.5 1.830 1.834 10.798 10.830 prepare_preconditioner 14 7.8 0.000 0.000 10.646 10.647 make_preconditioner 14 8.8 0.000 0.000 10.646 10.647 multiply_cannon 2077 13.5 0.323 0.324 9.268 9.279 init_scf_run 11 5.9 0.000 0.000 9.007 9.007 scf_env_initial_rho_setup 11 6.9 0.000 0.001 9.006 9.006 build_core_ppl_forces 11 5.9 8.783 8.910 8.783 8.910 pw_gpu_sf 1053 16.2 8.393 8.395 8.393 8.395 multiply_cannon_loop 2077 14.5 0.261 0.263 8.207 8.222 pw_gpu_fg 947 15.5 7.624 7.728 7.624 7.728 make_m2s 4154 13.5 0.047 0.048 7.699 7.705 ot_mini 106 10.8 0.001 0.001 7.633 7.637 build_core_hamiltonian_matrix 11 6.9 0.001 0.001 7.577 7.579 make_images 4154 14.5 1.009 1.018 7.514 7.517 wfi_extrapolate 11 7.9 0.002 0.002 7.041 7.041 pw_gpu_ffc 1053 16.2 6.036 6.069 6.036 6.069 make_full_inverse_cholesky 14 9.8 0.000 0.000 5.741 5.897 build_overlap_matrix_low 22 6.9 5.534 5.563 5.616 5.647 build_kinetic_matrix_low 22 6.9 5.381 5.413 5.475 5.508 qs_ot_get_derivative 106 11.8 0.002 0.002 4.806 4.811 pw_poisson_solve 117 10.5 0.003 0.003 4.732 4.734 pw_gpu_cff 947 15.5 4.618 4.637 4.618 4.637 transfer_rs2pw 479 10.8 0.010 0.010 4.422 4.571 make_full_single_inverse 14 9.8 0.002 0.002 4.307 4.307 multiply_cannon_multrec 4154 15.5 1.760 1.772 4.284 4.305 pw_derive 1053 13.8 4.072 4.087 4.072 4.087 make_images_data 4154 15.5 0.059 0.060 4.046 4.054 qs_env_update_s_mstruct 11 6.9 0.000 0.000 3.874 3.891 transfer_rs2pw_200 128 11.7 2.624 2.671 3.671 3.824 hybrid_alltoall_any 4154 16.5 2.806 2.809 3.805 3.812 copy_dbcsr_to_fm 143 10.8 0.089 0.090 3.605 3.616 mp_waitall_1 58635 17.0 3.492 3.525 3.492 3.525 build_core_ppl 11 7.9 3.418 3.453 3.418 3.453 transfer_dbcsr_to_fm 14 10.8 0.002 0.002 3.226 3.227 transfer_pw2rs 479 13.4 0.006 0.006 3.177 3.179 dbcsr_to_fm_plan_create 14 12.8 2.733 2.802 3.042 3.043 ot_diis_step 106 11.8 0.006 0.006 2.806 2.806 arnoldi_generalized_ev 14 10.8 0.000 0.000 2.704 2.705 fft_wrap_pw1pw2_70 234 13.2 0.002 0.002 2.678 2.698 pw_copy 1755 13.0 2.686 2.690 2.686 2.690 dbcsr_sym_matrix_vector_mult 1269 12.5 0.038 0.038 2.661 2.663 transfer_pw2rs_200 128 14.1 1.637 1.645 2.551 2.555 qs_create_task_list 11 7.9 0.000 0.000 2.501 2.508 generate_qs_task_list 11 8.9 1.403 1.407 2.501 2.507 gev_build_subspace 23 11.5 0.011 0.011 2.493 2.493 apply_preconditioner_dbcsr 120 12.8 0.000 0.000 2.396 2.398 apply_single 120 13.8 0.001 0.001 2.396 2.398 dbcsr_sym_matrix_vector_mult_l 1269 13.5 2.273 2.282 2.279 2.288 dbcsr_mm_accdrv_process 9444 16.2 0.840 1.049 2.248 2.252 qs_ot_get_derivative_taylor 89 12.9 0.004 0.004 2.224 2.227 calculate_dm_sparse 117 9.7 0.001 0.001 2.144 2.144 pw_poisson_set 118 11.5 0.005 0.005 2.122 2.124 cp_dbcsr_sm_fm_multiply 46 9.3 0.002 0.002 1.897 1.900 pw_integral_ab_c1d_c1d_gs 117 11.5 1.834 1.834 1.855 1.855 multiply_cannon_sync_h2d 4154 15.5 1.787 1.802 1.787 1.802 qs_ot_get_p 120 10.5 0.001 0.001 1.760 1.764 pw_axpy 1170 12.0 1.592 1.599 1.592 1.599 dbcsr_complete_redistribute 309 11.8 0.556 0.578 1.397 1.559 dbcsr_special_finalize 6231 15.5 0.036 0.037 1.520 1.525 cp_dbcsr_sm_fm_multiply_core 46 10.3 0.000 0.000 1.460 1.464 cp_fm_cholesky_invert 14 10.8 1.422 1.422 1.422 1.422 dbcsr_merge_single_wm 4154 16.5 0.138 0.139 1.401 1.405 mp_sendrecv_dv 479 12.8 1.284 1.386 1.284 1.386 copy_fm_to_dbcsr 180 10.8 0.002 0.002 1.186 1.346 multiply_cannon_metrocomm1 4154 15.5 0.013 0.014 1.276 1.329 calculate_rho_core 11 7.9 0.174 0.175 1.314 1.320 dbcsr_dot 1125 12.2 1.195 1.198 1.277 1.279 calculate_first_density_matrix 1 7.0 0.000 0.000 1.156 1.156 jit_kernel_multiply 12 15.0 0.893 1.101 0.893 1.101 dbcsr_sort_data 4154 17.5 0.975 0.979 0.975 0.979 cp_dbcsr_plus_fm_fm_t 22 8.9 0.001 0.001 0.943 0.943 qs_ot_get_orbitals 106 10.8 0.001 0.001 0.834 0.835 qs_ot_p2m_diag 19 11.0 0.036 0.036 0.830 0.830 dbcsr_copy 7924 13.4 0.206 0.207 0.826 0.827 build_core_ppnl_forces 11 5.9 0.794 0.813 0.794 0.813 evaluate_core_matrix_traces 117 8.5 0.001 0.001 0.803 0.805 calculate_ptrace_kp 234 9.5 0.001 0.001 0.802 0.804 grid_create_task_list 11 9.9 0.772 0.778 0.772 0.778 transfer_fm_to_dbcsr 14 9.8 0.000 0.000 0.598 0.757 fft_wrap_pw1pw2_30 234 13.2 0.001 0.001 0.709 0.722 cp_dbcsr_syevd 19 12.0 0.002 0.002 0.703 0.703 cp_fm_diag_elpa 19 13.0 0.000 0.000 0.674 0.674 cp_fm_diag_elpa_base 19 14.0 0.663 0.666 0.673 0.673 cp_fm_cholesky_decompose 28 10.5 0.670 0.672 0.670 0.672 make_images_pack 4154 15.5 0.645 0.645 0.660 0.660 qs_init_subsys 1 2.0 0.001 0.001 0.650 0.650 dbcsr_finalize 4569 14.0 0.062 0.062 0.638 0.647 qs_env_setup 1 3.0 0.000 0.000 0.642 0.642 qs_env_rebuild_pw_env 23 5.3 0.000 0.000 0.641 0.642 pw_env_rebuild 1 5.0 0.000 0.000 0.641 0.642 cp_fm_uplo_to_full 47 13.4 0.455 0.617 0.455 0.617 pw_grid_setup 4 6.0 0.000 0.000 0.615 0.616 pw_grid_setup_internal 4 7.0 0.007 0.007 0.605 0.606 transfer_rs2pw_70 117 11.9 0.397 0.398 0.582 0.585 qs_ot_get_derivative_diag 17 12.0 0.001 0.001 0.564 0.565 dbcsr_copy_into_existing 22 7.9 0.561 0.563 0.562 0.563 dbcsr_merge_all 4154 15.2 0.182 0.185 0.540 0.549 make_basis_sm 14 9.3 0.001 0.001 0.546 0.547 acc_transpose_blocks 4154 15.5 0.024 0.024 0.543 0.543 pw_zero 585 13.0 0.537 0.542 0.537 0.542 mp_sum_d 3821 11.6 0.426 0.522 0.426 0.522 dbcsr_mm_accdrv_process_sort 9444 17.2 0.515 0.519 0.515 0.519 mp_alltoall_d11v 1526 13.9 0.485 0.495 0.485 0.495 pw_grid_sort 4 8.0 0.365 0.366 0.492 0.493 mp_sum_l 6260 13.5 0.389 0.489 0.389 0.489 transfer_pw2rs_70 117 14.5 0.317 0.317 0.485 0.486 dbcsr_sort_indices 10944 16.6 0.438 0.441 0.438 0.441 parallel_gemm_fm_cosma 96 8.9 0.421 0.424 0.421 0.424 ot_scf_init 14 7.8 0.002 0.002 0.402 0.403 compute_matrix_w 11 5.9 0.000 0.000 0.399 0.400 calculate_w_matrix_ot 11 6.9 0.003 0.003 0.399 0.400 reorthogonalize_vectors 10 9.0 0.000 0.000 0.387 0.387 mp_alltoall_i22 504 14.0 0.213 0.363 0.213 0.363 cp_dbcsr_alloc_block_from_nbl 88 7.7 0.233 0.235 0.356 0.360 build_qs_neighbor_lists 11 6.9 0.001 0.001 0.350 0.359 integrate_v_core_rspace 11 7.9 0.070 0.071 0.338 0.341 dbcsr_add_d 1879 13.1 0.003 0.003 0.331 0.331 dbcsr_add_anytype 1879 14.1 0.174 0.177 0.328 0.328 distribute_tasks 11 9.9 0.310 0.312 0.310 0.312 pw_scale 468 12.0 0.301 0.301 0.301 0.301 calculate_ecore_overlap 22 5.9 0.001 0.001 0.180 0.297 setup_rec_index_2d 4154 14.5 0.290 0.293 0.290 0.293 multiply_cannon_multrec_finali 2077 16.5 0.005 0.005 0.277 0.281 mp_alltoall_i 14 13.8 0.207 0.280 0.207 0.280 dbcsr_mm_multrec_finalize 2077 17.5 0.022 0.023 0.272 0.276 fft_wrap_pw1pw2_10 234 13.2 0.001 0.001 0.266 0.269 dbcsr_make_untransposed_blocks 2537 13.5 0.243 0.243 0.255 0.256 pw_multiply_with 117 11.5 0.254 0.255 0.254 0.255 acc_transpose_blocks_sync 12462 16.5 0.253 0.254 0.253 0.254 dbcsr_mm_sched_finalize 2077 18.5 0.244 0.247 0.249 0.252 build_core_ppnl 11 7.9 0.241 0.248 0.241 0.248 acc_transpose_blocks_kernels 4154 16.5 0.054 0.055 0.241 0.243 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="w64PBE", label="w64PBE", y=242.456, yerr=0.0 Plot: name="w64PBE_timings_6cpu_1gpu", title="Timings of w64PBE with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="rest", label="rest", y=125.34599999999999, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=41.901, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=33.193, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=19.057, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="pbe_lda_eval", label="pbe_lda_eval", y=12.212, yerr=0.0 PlotPoint: plot="w64PBE_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=10.747, yerr=0.0 Running w64SCAN.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/w64SCAN_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.242 0.246 934.173 934.173 qs_mol_dyn_low 1 2.0 0.004 0.004 931.785 931.788 qs_forces 11 3.9 0.002 0.003 931.731 931.733 qs_energies 11 4.9 0.002 0.002 838.872 838.872 scf_env_do_scf 11 5.9 0.001 0.002 798.944 798.944 velocity_verlet 10 3.0 0.002 0.002 740.386 740.405 scf_env_do_scf_inner_loop 106 6.8 0.006 0.009 716.631 716.631 rebuild_ks_matrix 117 8.5 0.001 0.001 639.884 639.886 qs_ks_build_kohn_sham_matrix 117 9.5 0.028 0.028 639.883 639.885 qs_ks_update_qs_env 119 7.8 0.002 0.002 561.883 561.886 fft_wrap_pw1pw2 3053 12.6 0.077 0.077 455.319 456.003 fft_wrap_pw1pw2_400 1649 13.9 0.012 0.012 437.081 437.718 qs_vxc_create 117 10.5 0.003 0.003 393.770 393.781 xc_vxc_pw_create 117 11.5 4.768 4.778 393.768 393.779 xc_rho_set_and_dset_create 117 12.5 6.291 6.293 257.591 258.166 qs_rho_update_rho_low 117 7.9 0.002 0.002 233.679 233.686 calculate_rho_elec 234 8.9 6.948 6.956 233.677 233.684 pw_gpu_r3dc1d_3d_ps 1532 14.1 132.554 132.573 227.563 228.254 pw_gpu_c1dr3d_3d_ps 1521 15.1 130.544 130.552 227.660 227.667 xc_pw_derive 702 13.5 0.012 0.012 196.053 196.608 sum_up_and_integrate 117 10.5 0.008 0.008 190.613 190.995 integrate_v_rspace 234 11.5 0.438 0.439 189.724 190.110 density_rs2pw 234 9.9 0.023 0.023 172.331 172.751 xc_functional_eval 234 13.5 0.004 0.004 150.184 150.755 libxc_spin_unpolarized_eval 234 14.5 149.733 150.304 150.179 150.751 xc_pw_divergence 117 12.5 0.007 0.007 130.012 130.561 potential_pw2rs 234 12.5 0.292 0.295 101.864 102.068 grid_integrate_task_list 234 12.5 87.420 88.010 87.420 88.010 init_scf_loop 13 6.8 0.001 0.001 82.260 82.260 qs_ks_update_qs_env_forces 11 4.9 0.000 0.000 78.803 78.804 mp_alltoall_z22v 3053 16.6 74.773 75.383 74.773 75.383 grid_collocate_task_list 234 9.9 54.260 54.693 54.260 54.693 yz_to_x 1532 15.1 7.901 7.911 46.489 47.140 x_to_yz 1521 16.1 9.255 9.258 45.440 45.472 transfer_rs2pw 947 10.9 0.022 0.022 37.034 37.546 transfer_rs2pw_400 245 11.8 26.352 26.453 32.440 32.975 pw_gpu_sf 1521 16.1 31.384 31.401 31.384 31.401 pw_gpu_fg 1532 15.1 30.858 30.903 30.858 30.903 transfer_pw2rs 947 13.5 0.017 0.017 30.587 30.591 transfer_pw2rs_400 245 14.3 21.724 21.733 27.236 27.238 init_scf_run 11 5.9 0.000 0.000 25.579 25.579 scf_env_initial_rho_setup 11 6.9 0.000 0.001 25.579 25.579 wfi_extrapolate 11 7.9 0.002 0.002 21.904 21.904 pw_gpu_ffc 1521 16.1 20.263 20.281 20.263 20.281 dbcsr_multiply_generic 2139 12.6 0.153 0.154 19.520 19.889 pw_poisson_solve 117 10.5 0.004 0.004 17.821 17.821 pw_gpu_cff 1532 15.1 17.506 17.520 17.506 17.520 fft_wrap_pw1pw2_140 468 13.2 0.003 0.003 14.346 14.461 qs_scf_new_mos 106 7.8 0.001 0.001 14.162 14.167 qs_scf_loop_do_ot 106 8.8 0.001 0.001 14.161 14.167 build_core_hamiltonian_matrix_ 11 4.9 0.001 0.001 13.812 14.042 qs_energies_init_hamiltonians 11 5.9 0.000 0.000 13.842 13.842 ot_scf_mini 106 9.8 0.003 0.003 12.693 12.693 pw_derive 1053 13.8 12.395 12.424 12.395 12.424 prepare_preconditioner 13 7.8 0.000 0.000 11.049 11.052 make_preconditioner 13 8.8 0.000 0.000 11.048 11.052 multiply_cannon 2139 13.6 0.325 0.331 9.682 9.696 pw_copy 2223 13.1 9.240 9.251 9.240 9.251 mp_waitall_1 60839 17.0 9.136 9.212 9.136 9.212 pw_integral_ab_c1d_c1d_gs 117 11.5 8.313 8.334 8.697 8.698 multiply_cannon_loop 2139 14.6 0.268 0.268 8.596 8.602 make_m2s 4278 13.6 0.047 0.048 7.805 7.813 ot_mini 106 10.8 0.001 0.001 7.771 7.772 mp_sendrecv_dv 947 12.9 7.359 7.752 7.359 7.752 make_images 4278 14.6 1.039 1.039 7.615 7.624 qs_env_update_s_mstruct 11 6.9 0.000 0.000 7.263 7.281 pw_poisson_set 118 11.5 0.006 0.007 6.757 6.757 build_core_ppl_forces 11 5.9 6.393 6.623 6.393 6.623 build_core_hamiltonian_matrix 11 6.9 0.001 0.001 6.215 6.287 pw_axpy 1638 11.7 6.127 6.135 6.127 6.135 make_full_inverse_cholesky 13 9.8 0.000 0.000 5.583 5.728 calculate_rho_core 11 7.9 0.446 0.446 5.168 5.241 qs_ot_get_derivative 106 11.8 0.002 0.002 4.939 4.940 make_full_single_inverse 13 9.8 0.002 0.002 4.877 4.878 build_overlap_matrix_low 22 6.9 4.757 4.759 4.833 4.837 build_kinetic_matrix_low 22 6.9 4.572 4.575 4.658 4.661 multiply_cannon_multrec 4278 15.6 1.818 1.820 4.610 4.628 make_images_data 4278 15.6 0.058 0.058 4.068 4.085 transfer_rs2pw_140 234 11.9 2.838 2.854 3.859 3.876 hybrid_alltoall_any 4278 16.6 2.799 2.824 3.828 3.847 copy_dbcsr_to_fm 138 10.8 0.087 0.087 3.634 3.638 transfer_dbcsr_to_fm 13 10.8 0.003 0.005 3.248 3.248 arnoldi_generalized_ev 13 10.8 0.000 0.000 3.142 3.142 dbcsr_sym_matrix_vector_mult 1206 12.5 0.039 0.039 3.096 3.098 dbcsr_to_fm_plan_create 13 12.8 2.671 2.833 3.051 3.051 fft_wrap_pw1pw2_50 468 13.2 0.003 0.003 2.912 2.978 gev_build_subspace 22 11.5 0.012 0.012 2.922 2.922 ot_diis_step 106 11.8 0.006 0.006 2.810 2.810 transfer_pw2rs_140 234 14.5 1.719 1.720 2.703 2.705 dbcsr_sym_matrix_vector_mult_l 1206 13.5 2.494 2.690 2.501 2.697 build_core_ppl 11 7.9 2.603 2.646 2.603 2.646 dbcsr_mm_accdrv_process 9536 16.3 0.625 0.891 2.507 2.542 apply_preconditioner_dbcsr 119 12.8 0.000 0.000 2.367 2.369 apply_single 119 13.8 0.001 0.001 2.367 2.369 qs_ot_get_derivative_taylor 89 12.9 0.004 0.004 2.333 2.334 calculate_dm_sparse 117 9.7 0.001 0.001 2.196 2.202 pw_zero 702 12.6 2.185 2.190 2.185 2.190 qs_init_subsys 1 2.0 0.001 0.001 2.054 2.054 qs_env_setup 1 3.0 0.000 0.000 2.046 2.047 qs_env_rebuild_pw_env 23 5.3 0.000 0.000 2.046 2.047 pw_env_rebuild 1 5.0 0.000 0.000 2.045 2.046 cp_dbcsr_sm_fm_multiply 45 9.4 0.002 0.002 2.026 2.030 pw_grid_setup 4 6.0 0.000 0.000 1.981 1.982 pw_grid_setup_internal 4 7.0 0.021 0.021 1.949 1.950 qs_create_task_list 11 7.9 0.000 0.000 1.868 1.926 generate_qs_task_list 11 8.9 0.950 0.954 1.868 1.925 multiply_cannon_sync_h2d 4278 15.6 1.775 1.822 1.775 1.822 qs_ot_get_p 119 10.6 0.001 0.001 1.804 1.810 pw_grid_sort 4 8.0 1.193 1.196 1.610 1.611 mp_sum_d 3883 11.6 1.303 1.596 1.303 1.596 jit_kernel_multiply 13 15.2 1.349 1.578 1.349 1.578 cp_dbcsr_sm_fm_multiply_core 45 10.4 0.000 0.000 1.564 1.564 dbcsr_special_finalize 6417 15.6 0.037 0.038 1.541 1.545 dbcsr_complete_redistribute 299 11.7 0.568 0.578 1.399 1.542 dbcsr_merge_single_wm 4278 16.6 0.142 0.147 1.419 1.422 multiply_cannon_metrocomm1 4278 15.6 0.013 0.013 1.312 1.382 integrate_v_core_rspace 11 7.9 0.156 0.156 1.366 1.366 copy_fm_to_dbcsr 174 10.8 0.002 0.002 1.177 1.322 cp_fm_cholesky_invert 13 10.8 1.312 1.313 1.312 1.313 dbcsr_dot 1134 12.2 1.197 1.208 1.284 1.285 calculate_first_density_matrix 1 7.0 0.000 0.000 1.188 1.188 mp_sum_l 6446 13.6 0.802 1.172 0.802 1.172 pw_scale 585 11.9 1.151 1.156 1.151 1.156 dbcsr_sort_data 4278 17.6 0.985 0.985 0.985 0.985 cp_dbcsr_plus_fm_fm_t 22 8.9 0.001 0.001 0.976 0.978 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="w64SCAN", label="w64SCAN", y=934.173, yerr=0.0 Plot: name="w64SCAN_timings_6cpu_1gpu", title="Timings of w64SCAN with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="rest", label="rest", y=359.149, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="libxc_spin_unpolarized_eval", label="libxc_spin_unpolarized_eval", y=149.733, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="pw_gpu_r3dc1d_3d_ps", label="pw_gpu_r3dc1d_3d_ps", y=132.554, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=130.544, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=87.42, yerr=0.0 PlotPoint: plot="w64SCAN_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=74.773, yerr=0.0 Running ZnO.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/ZnO_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.025 0.027 87.231 87.231 qs_energies 1 2.0 0.001 0.001 85.833 85.833 scf_env_do_scf 1 3.0 0.000 0.000 84.801 84.801 scf_env_do_scf_inner_loop 10 4.0 0.004 0.005 84.801 84.801 qs_scf_new_mos_kp 10 5.0 0.000 0.000 82.277 82.352 do_general_diag_kp 10 6.0 33.167 33.177 82.277 82.352 cp_cfm_geeig_local 15087 7.0 30.662 30.917 30.662 30.917 kpoint_density_transform 10 7.0 0.015 0.016 16.873 16.883 kpoint_density_transform_regul 10 8.0 0.152 0.157 16.845 16.845 kp_density_fft 10 9.0 2.014 2.046 13.336 13.337 k_grid_to_cell_fft 910 10.0 2.595 2.628 9.377 9.478 fft3d_s 28881 11.0 6.757 6.824 6.783 6.850 mp_alltoall_z11v 910 10.0 1.945 2.076 1.945 2.076 qs_ks_update_qs_env 10 5.0 0.000 0.000 1.848 1.924 rebuild_ks_matrix 10 6.0 0.000 0.000 1.699 1.775 qs_ks_build_kohn_sham_matrix 10 7.0 0.009 0.009 1.699 1.775 copy_fm_to_dbcsr 8860 9.0 0.049 0.049 1.754 1.768 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="ZnO", label="ZnO", y=87.231, yerr=0.0 Plot: name="ZnO_timings_6cpu_1gpu", title="Timings of ZnO with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="rest", label="rest", y=12.036000000000001, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="do_general_diag_kp", label="do_general_diag_kp", y=33.167, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="cp_cfm_geeig_local", label="cp_cfm_geeig_local", y=30.662, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="fft3d_s", label="fft3d_s", y=6.757, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="k_grid_to_cell_fft", label="k_grid_to_cell_fft", y=2.595, yerr=0.0 PlotPoint: plot="ZnO_timings_6cpu_1gpu", name="kp_density_fft", label="kp_density_fft", y=2.014, yerr=0.0 Running GW_PBE_4benzene.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/GW_PBE_4benzene_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.021 0.024 107.396 107.397 qs_energies 1 2.0 0.000 0.000 107.058 107.064 mp2_main 1 3.0 0.000 0.000 100.022 100.028 mp2_gpw_main 1 4.0 0.000 0.000 98.314 98.321 rpa_ri_compute_en 1 5.0 0.000 0.000 91.243 91.249 rpa_num_int 1 6.0 0.001 0.001 91.233 91.239 dbt_total 2336 9.6 0.022 0.022 73.012 73.013 compute_mat_P_omega 1 7.0 0.002 0.002 69.640 69.650 compute_mat_P_omega_contract 10 8.0 5.317 5.348 69.301 69.316 dbt_contract 787 11.0 0.051 0.051 47.219 47.220 dbt_tas_total 1149 12.2 0.149 0.150 36.298 36.299 dbt_tas_multiply 807 12.1 0.003 0.003 35.599 35.599 dbt_tas_dbm 807 14.1 0.006 0.006 27.123 27.123 dbm_multiply 807 16.1 25.803 26.612 25.803 26.612 dbt_copy 1107 10.7 0.068 0.068 26.048 26.342 compute_mat_P_omega_calc_M_occ 250 9.0 5.354 5.394 24.115 24.115 dbt_tas_mm_1N 524 15.1 0.003 0.003 17.131 17.928 dbt_reshape 594 11.8 7.142 7.441 17.179 17.292 compute_QP_energies 1 7.0 0.000 0.000 15.436 15.436 compute_self_energy_cubic_gw 1 8.0 0.118 0.120 15.436 15.436 compute_mat_P_omega_calc_M_vir 250 9.0 0.001 0.001 15.067 15.067 dbt_tas_reserve_blocks_index 3266 14.3 0.697 0.711 11.118 11.269 dbm_reserve_blocks 3634 15.3 10.733 10.902 10.733 10.902 dbt_crop 1042 12.0 6.979 7.114 9.307 9.515 dbt_reserve_blocks_index 2347 13.0 0.337 0.346 9.378 9.407 dbt_reserve_blocks_index_array 2289 12.1 0.011 0.011 9.174 9.221 compute_mat_P_omega_calc_P_t 250 9.0 0.001 0.001 9.009 9.009 mp_waitall_2 2656 15.9 8.218 8.316 8.218 8.316 dbt_tas_mm_2 251 15.0 0.003 0.003 7.670 7.670 dbt_communicate_buffer 594 12.8 0.013 0.014 7.398 7.489 mp2_ri_gpw_compute_in 1 5.0 0.001 0.001 7.061 7.061 contract_cubic_gw 21 9.0 0.000 0.000 6.870 6.870 scf_env_do_scf 1 3.0 0.000 0.000 6.460 6.460 scf_env_do_scf_inner_loop 17 4.0 0.001 0.001 6.460 6.460 compute_mat_P_omega_copy_M_vir 250 9.0 0.002 0.002 5.590 5.605 compute_mat_P_omega_copy_M_occ 250 9.0 0.002 0.002 5.458 5.462 dbt_tas_copy 511 11.5 2.546 2.624 4.423 4.609 dbcsr_multiply_generic 30 8.1 0.002 0.003 4.535 4.582 multiply_cannon 30 9.1 0.009 0.013 4.339 4.382 multiply_cannon_loop 30 10.1 0.004 0.004 4.282 4.327 multiply_cannon_multrec 60 11.1 0.251 0.258 3.737 3.755 qs_scf_new_mos 17 5.0 0.001 0.001 3.510 3.550 trace_sigma_gw 21 9.0 0.559 0.584 3.522 3.522 mp_sync 8688 11.6 2.889 3.459 2.889 3.459 dbcsr_mm_accdrv_process 328 12.3 0.022 0.022 3.204 3.218 jit_kernel_multiply 17 11.6 3.176 3.189 3.176 3.189 dbt_split_copyback 70 10.6 1.252 1.296 2.879 2.917 get_2c_integrals 1 6.0 0.000 0.000 2.688 2.688 fft_wrap_pw1pw2 301 10.2 0.005 0.006 2.535 2.543 convert_to_new_pgrid 2421 14.1 0.038 0.039 2.447 2.452 mp2_ri_gpw_compute_in_copy_3c 6 6.0 0.224 0.226 2.275 2.417 dbm_copy 1614 15.1 2.410 2.416 2.410 2.416 qs_ks_build_kohn_sham_matrix 18 6.9 0.003 0.003 2.409 2.410 qs_ks_update_qs_env 17 5.0 0.000 0.000 2.377 2.379 rebuild_ks_matrix 17 6.0 0.000 0.000 2.370 2.371 build_3c_integrals 5 6.0 1.457 1.491 2.090 2.231 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="GW_PBE_4benzene", label="GW_PBE_4benzene", y=107.396, yerr=0.0 Plot: name="GW_PBE_4benzene_timings_6cpu_1gpu", title="Timings of GW_PBE_4benzene with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="rest", label="rest", y=48.521, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbm_multiply", label="dbm_multiply", y=25.803, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbm_reserve_blocks", label="dbm_reserve_blocks", y=10.733, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="mp_waitall_2", label="mp_waitall_2", y=8.218, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbt_reshape", label="dbt_reshape", y=7.142, yerr=0.0 PlotPoint: plot="GW_PBE_4benzene_timings_6cpu_1gpu", name="dbt_crop", label="dbt_crop", y=6.979, yerr=0.0 Running RI-HFX_H2O-32.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/RI-HFX_H2O-32_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.031 0.035 214.723 214.724 qs_forces 1 2.0 0.000 0.000 214.226 214.226 rebuild_ks_matrix 7 6.6 0.000 0.000 210.041 210.041 qs_ks_build_kohn_sham_matrix 7 7.6 0.002 0.002 210.041 210.041 hfx_ks_matrix 7 8.6 0.000 0.000 206.127 206.128 dbt_total 849 11.0 0.010 0.010 153.591 153.591 hfx_ri_update_ks 7 9.6 0.000 0.000 119.964 119.964 hfx_ri_update_ks_Pmat 7 10.6 22.863 22.890 119.959 119.959 qs_energies 1 3.0 0.000 0.000 115.143 115.143 scf_env_do_scf 1 4.0 0.000 0.000 113.133 113.133 qs_ks_update_qs_env 8 6.0 0.000 0.000 111.004 111.004 qs_ks_update_qs_env_forces 1 3.0 0.000 0.000 99.044 99.044 dbt_contract 207 12.4 0.055 0.056 87.887 87.888 hfx_ri_update_forces 1 7.0 1.130 1.175 86.161 86.162 dbt_tas_total 369 13.4 0.086 0.087 70.723 70.723 dbt_tas_multiply 216 13.5 0.001 0.001 67.667 67.668 dbt_copy 423 11.8 0.047 0.047 60.429 60.729 scf_env_do_scf_inner_loop 6 5.0 0.000 0.001 59.322 59.322 init_scf_loop 2 5.0 0.000 0.000 53.810 53.810 dbt_tas_dbm 216 15.5 0.002 0.002 52.832 52.832 dbm_multiply 216 17.5 49.473 49.811 49.473 49.811 hfx_ri_forces_Pmat_3c 1 8.0 3.494 3.498 48.281 48.314 dbt_reshape 175 13.2 20.642 20.820 46.037 46.096 hfx_ri_update_ks_Pmat_KS 63 11.6 0.001 0.001 34.120 34.120 precalc_derivatives 1 8.0 2.037 2.107 30.998 30.999 mp_waitall_2 1022 16.5 23.785 23.867 23.785 23.867 dbt_tas_mm_2 91 16.5 0.001 0.001 22.569 22.569 dbt_crop 372 13.7 15.175 15.214 19.595 19.722 dbt_communicate_buffer 175 14.2 0.005 0.005 19.600 19.652 dbt_tas_reserve_blocks_index 1323 15.4 1.861 1.873 19.216 19.506 hfx_ri_pre_scf_Pmat 1 12.0 0.000 0.000 19.120 19.120 dbm_reserve_blocks 1491 16.3 18.013 18.316 18.013 18.316 hfx_ri_update_ks_Pmat_copy_2 63 11.6 0.000 0.000 17.641 17.641 build_3c_derivatives 3 9.0 3.053 3.236 17.025 17.053 hfx_ri_update_ks_Pmat_Px3C 63 11.6 0.000 0.000 16.619 16.620 dbt_tas_mm_3T 77 17.1 0.001 0.001 16.347 16.510 dbt_reserve_blocks_index 889 14.5 0.669 0.674 15.451 15.615 dbt_reserve_blocks_index_array 859 13.5 0.008 0.008 15.153 15.302 dbt_tas_mm_3N 37 15.4 0.000 0.000 11.238 11.242 dbt_tas_copy 248 12.5 4.729 4.731 8.917 9.043 mp_sync 2901 12.8 7.700 8.977 7.700 8.977 hfx_ri_pre_scf_Pmat_copy_2 9 13.0 2.113 2.143 5.553 5.583 hfx_ri_pre_scf_Pmat_int 1 13.0 0.000 0.000 5.551 5.551 dbt_tas_replicate 168 15.1 2.398 2.423 5.367 5.406 hfx_ri_pre_scf_calc_tensors 1 14.0 0.003 0.003 4.781 4.799 hfx_ri_pre_scf_Pmat_RIx3C 9 13.0 0.000 0.000 4.432 4.483 dbt_tas_reserve_blocks_templat 266 13.6 0.111 0.117 4.250 4.376 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="RI-HFX_H2O-32", label="RI-HFX_H2O-32", y=214.723, yerr=0.0 Plot: name="RI-HFX_H2O-32_timings_6cpu_1gpu", title="Timings of RI-HFX_H2O-32 with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="rest", label="rest", y=79.947, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbm_multiply", label="dbm_multiply", y=49.473, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="mp_waitall_2", label="mp_waitall_2", y=23.785, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="hfx_ri_update_ks_Pmat", label="hfx_ri_update_ks_Pmat", y=22.863, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbt_reshape", label="dbt_reshape", y=20.642, yerr=0.0 PlotPoint: plot="RI-HFX_H2O-32_timings_6cpu_1gpu", name="dbm_reserve_blocks", label="dbm_reserve_blocks", y=18.013, yerr=0.0 Running RI-MP2_ammonia.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/RI-MP2_ammonia_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.012 0.014 105.495 105.495 qs_energies 1 2.0 0.000 0.000 105.291 105.291 mp2_main 1 3.0 0.000 0.000 96.579 96.579 mp2_gpw_main 1 4.0 0.001 0.002 96.018 96.018 mp2_ri_gpw_compute_in 1 5.0 0.557 0.560 51.619 51.638 mp2_ri_gpw_compute_en 1 5.0 0.108 0.112 44.332 44.351 mp2_ri_gpw_compute_in_loop 1 6.0 0.014 0.015 43.211 43.227 mp2_ri_gpw_compute_en_RI_loop 1 6.0 12.976 13.013 41.624 41.625 dbcsr_multiply_generic 2666 8.0 0.169 0.169 22.511 22.576 ao_to_mo_and_store_B_mult_1 1328 7.0 0.014 0.014 21.593 21.658 mp2_eri_3c_integrate_gpw 1328 7.0 0.019 0.019 16.368 16.524 mp2_ri_gpw_compute_en_expansio 1040 7.0 0.742 0.742 16.406 16.483 local_gemm 1040 8.0 15.664 15.742 15.664 15.742 make_m2s 5332 9.0 0.057 0.059 12.742 12.805 make_images 5332 10.0 2.258 2.274 12.560 12.630 multiply_cannon 2666 9.0 0.414 0.416 9.074 9.197 make_images_data 5332 11.0 0.072 0.072 8.589 8.648 hybrid_alltoall_any 5332 12.0 8.381 8.442 8.409 8.469 fft_wrap_pw1pw2 26668 10.4 0.145 0.146 7.929 8.178 multiply_cannon_loop 2666 10.0 0.205 0.209 7.931 8.047 integrate_v_rspace 1338 8.0 1.057 1.062 7.844 7.865 get_2c_integrals 1 6.0 0.004 0.005 7.849 7.849 scf_env_do_scf 1 3.0 0.000 0.000 7.770 7.771 scf_env_do_scf_inner_loop 10 4.0 0.001 0.001 7.770 7.771 collocate_function 1328 8.0 5.209 5.256 7.355 7.517 compute_2c_integrals 1 7.0 0.007 0.007 7.277 7.278 compute_2c_integrals_loop_lm 1 8.0 0.022 0.023 7.025 7.110 mp2_eri_2c_integrate_gpw 1 9.0 2.081 2.110 7.003 7.088 mp2_ri_gpw_compute_en_comm 221 7.0 1.055 1.063 5.959 6.111 qs_scf_new_mos 10 5.0 0.000 0.000 6.065 6.068 grid_integrate_task_list 1338 9.0 5.441 5.454 5.441 5.454 mp2_ri_gpw_compute_en_ener 1040 7.0 5.116 5.155 5.116 5.155 ao_to_mo_and_store_B_E_Ex_1 1328 7.0 3.490 3.534 4.991 5.054 fft_wrap_pw1pw2_20 10647 11.4 0.023 0.023 4.553 4.802 pw_gpu_r3dc1d_3d 13282 12.2 3.976 4.240 3.976 4.240 multiply_cannon_multrec 2676 11.0 1.929 2.032 4.045 4.159 mp_sendrecv_dm3 442 8.0 3.861 4.003 3.861 4.003 copy_dbcsr_to_fm 1351 8.0 0.092 0.094 3.467 3.482 eigensolver 11 5.8 0.001 0.002 3.401 3.405 potential_pw2rs 2666 10.0 0.104 0.104 2.767 2.779 cp_fm_diag_elpa 11 6.8 0.000 0.000 2.714 2.715 cp_fm_diag_elpa_base 11 7.8 2.627 2.645 2.713 2.713 pw_gpu_c1dr3d_3d 13280 12.7 2.691 2.702 2.691 2.702 fft_wrap_pw1pw2_10 15957 11.5 0.021 0.021 2.420 2.427 collocate_single_gaussian 1328 10.0 0.099 0.101 2.359 2.411 replicate_iaK_2intgroup 1 6.0 2.150 2.155 2.293 2.296 mp2_eri_2c_integrate_gpw_pot_l 1328 10.0 0.004 0.004 2.225 2.228 fill_local_i_aL 884 7.5 2.210 2.212 2.210 2.212 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="RI-MP2_ammonia", label="RI-MP2_ammonia", y=105.495, yerr=0.0 Plot: name="RI-MP2_ammonia_timings_6cpu_1gpu", title="Timings of RI-MP2_ammonia with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="rest", label="rest", y=57.824000000000005, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="local_gemm", label="local_gemm", y=15.664, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="mp2_ri_gpw_compute_en_RI_loop", label="mp2_ri_gpw_compute_en_RI_loop", y=12.976, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="hybrid_alltoall_any", label="hybrid_alltoall_any", y=8.381, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=5.441, yerr=0.0 PlotPoint: plot="RI-MP2_ammonia_timings_6cpu_1gpu", name="collocate_function", label="collocate_function", y=5.209, yerr=0.0 Running diag_cu144_broy.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/diag_cu144_broy_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.096 0.103 197.471 197.471 qs_energies 1 2.0 0.000 0.000 196.324 196.325 scf_env_do_scf 1 3.0 0.000 0.000 182.104 182.104 scf_env_do_scf_inner_loop 15 4.0 0.001 0.002 182.104 182.104 qs_ks_update_qs_env 15 5.0 0.000 0.000 86.021 86.077 rebuild_ks_matrix 15 6.0 0.000 0.000 85.814 85.869 qs_ks_build_kohn_sham_matrix 15 7.0 0.003 0.004 85.814 85.869 qs_scf_new_mos 15 5.0 0.118 0.119 64.102 64.151 fft_wrap_pw1pw2 1071 10.6 0.029 0.029 54.834 54.879 eigensolver 15 6.0 0.002 0.002 45.845 45.921 qs_vxc_create 15 8.0 0.163 0.176 42.457 42.476 sum_up_and_integrate 15 8.0 0.001 0.001 41.816 41.897 integrate_v_rspace 15 9.0 0.047 0.048 41.791 41.871 grid_integrate_task_list 15 10.0 34.290 34.345 34.290 34.345 qs_rho_update_rho_low 16 5.0 0.000 0.000 29.457 29.457 calculate_rho_elec 16 6.0 0.185 0.186 29.457 29.457 cp_fm_diag_elpa 15 7.0 0.000 0.000 27.932 27.938 cp_fm_diag_elpa_base 15 8.0 26.113 26.678 27.926 27.927 fft_wrap_pw1pw2_150 735 11.9 0.005 0.005 27.728 27.759 calculate_vxc_nlvdw 15 9.0 1.086 1.100 27.630 27.635 pw_gpu_c1dr3d_3d_ps 555 12.7 5.673 5.776 27.458 27.514 pw_gpu_r3dc1d_3d_ps 516 12.6 6.241 6.263 27.340 27.351 grid_collocate_task_list 16 7.0 17.445 17.462 17.445 17.462 cp_fm_cholesky_restore 45 7.0 16.006 16.721 16.006 16.721 copy_dbcsr_to_fm 16 5.9 0.497 0.513 15.618 15.688 fft_wrap_pw1pw2_200 212 11.3 0.001 0.001 14.746 14.806 dbcsr_to_fm_plan_create 16 6.9 13.658 13.774 14.725 14.781 vdW_theta_forward 15 10.0 0.610 0.623 12.383 12.402 density_rs2pw 16 7.0 0.002 0.002 11.816 11.842 qs_energies_init_hamiltonians 1 3.0 0.000 0.000 10.253 10.253 vdW_theta_inverse 15 10.0 0.414 0.423 9.962 9.965 mp_alltoall_z22v 1071 14.6 9.501 9.652 9.501 9.652 pw_gpu_ffc 555 13.7 9.076 9.132 9.076 9.132 build_core_hamiltonian_matrix 1 4.0 0.000 0.000 8.860 8.982 pw_gpu_cff 516 13.6 8.669 8.675 8.669 8.675 xc_vxc_pw_create 15 9.0 0.186 0.188 7.921 7.928 potential_pw2rs 15 10.0 0.007 0.007 7.454 7.479 pw_gpu_fg 516 13.6 6.912 6.929 6.912 6.929 pw_gpu_sf 555 13.7 6.757 6.760 6.757 6.760 x_to_yz 555 13.7 0.974 0.980 5.920 6.020 fft_wrap_pw1pw2_10 62 10.5 0.000 0.000 5.538 5.540 yz_to_x 516 13.6 0.906 0.907 5.461 5.505 xc_pw_derive 90 11.0 0.001 0.001 5.061 5.097 prepare_nlvdw_density 15 9.0 0.069 0.070 5.084 5.088 build_core_ppnl 1 5.0 4.950 5.029 4.950 5.029 cp_fm_uplo_to_full 30 8.0 3.717 4.915 3.717 4.915 xc_rho_set_and_dset_create 15 10.0 0.129 0.134 4.318 4.345 gspace_mixing 14 5.0 0.129 0.130 4.151 4.151 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="diag_cu144_broy", label="diag_cu144_broy", y=197.471, yerr=0.0 Plot: name="diag_cu144_broy_timings_6cpu_1gpu", title="Timings of diag_cu144_broy with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="rest", label="rest", y=89.959, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="grid_integrate_task_list", label="grid_integrate_task_list", y=34.29, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="cp_fm_diag_elpa_base", label="cp_fm_diag_elpa_base", y=26.113, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="grid_collocate_task_list", label="grid_collocate_task_list", y=17.445, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="cp_fm_cholesky_restore", label="cp_fm_cholesky_restore", y=16.006, yerr=0.0 PlotPoint: plot="diag_cu144_broy_timings_6cpu_1gpu", name="dbcsr_to_fm_plan_create", label="dbcsr_to_fm_plan_create", y=13.658, yerr=0.0 Running bench_dftb.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/bench_dftb_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 1.995 2.014 162.373 162.373 qs_energies 1 2.0 0.000 0.000 160.267 160.267 ls_scf 1 3.0 0.000 0.000 153.102 153.103 ls_scf_main 1 4.0 0.001 0.001 141.569 141.569 density_matrix_trs4 5 5.0 0.004 0.004 112.859 112.910 dbcsr_multiply_generic 95 6.2 0.163 0.164 97.434 97.439 multiply_cannon 95 7.2 1.778 1.907 68.294 68.511 multiply_cannon_loop 95 8.2 0.173 0.173 57.137 57.386 multiply_cannon_multrec 190 9.2 43.368 43.880 48.637 49.163 ls_scf_dm_to_ks 5 5.0 0.000 0.000 26.834 26.883 make_m2s 190 7.2 0.016 0.016 24.519 24.533 make_images 190 8.2 5.411 5.593 23.974 23.987 matrix_ls_to_qs 5 6.0 0.000 0.000 17.772 17.960 dbcsr_complete_redistribute 11 7.5 10.776 10.971 15.198 15.353 matrix_decluster 5 7.0 0.000 0.000 13.886 14.040 arnoldi_extremal 6 6.2 0.000 0.000 11.739 11.743 arnoldi_normal_ev 6 7.2 0.005 0.005 11.739 11.743 build_subspace 12 8.2 0.033 0.034 11.499 11.500 qs_ks_update_qs_env 6 6.2 0.000 0.000 10.893 11.130 dbcsr_matrix_vector_mult 310 9.0 0.078 0.079 10.427 10.557 rebuild_ks_matrix 6 7.2 0.000 0.000 10.397 10.399 build_dftb_ks_matrix 6 8.2 0.001 0.001 10.397 10.399 make_images_data 190 9.2 0.007 0.007 10.102 10.158 build_dftb_coulomb 6 9.2 0.817 0.819 10.082 10.083 dbcsr_matrix_vector_mult_local 310 10.0 9.902 10.030 9.907 10.035 hybrid_alltoall_any 201 10.0 6.681 6.681 9.702 9.764 ls_scf_init_scf 1 4.0 0.000 0.000 9.729 9.730 tb_ewald_overlap 6 10.2 8.988 9.090 8.988 9.090 dbcsr_finalize 277 7.6 0.108 0.112 7.728 8.042 calculate_norms 380 9.2 7.800 8.013 7.800 8.013 ls_scf_init_matrix_S 1 5.0 0.000 0.000 7.860 7.865 dbcsr_merge_all 247 8.6 1.530 1.767 7.098 7.393 matrix_sqrt_Newton_Schulz 1 6.0 0.000 0.000 7.114 7.117 qs_energies_init_hamiltonians 1 3.0 0.000 0.000 7.105 7.106 build_qs_neighbor_lists 1 4.0 0.000 0.000 6.507 6.561 build_neighbor_lists_sab_tbe 1 5.0 6.322 6.377 6.322 6.377 dbcsr_copy 443 8.0 0.972 0.987 4.870 4.919 setup_rec_index_2d 190 8.2 4.829 4.856 4.829 4.856 dbcsr_special_finalize 285 9.2 0.005 0.005 4.818 4.824 dbcsr_dot 66 6.3 3.880 3.887 4.339 4.760 dbcsr_add_d 130 6.0 0.001 0.001 4.439 4.694 dbcsr_add_anytype 130 7.0 1.841 1.848 4.438 4.694 dbcsr_data_new 3509 9.3 4.459 4.589 4.459 4.589 dbcsr_sort_indices 643 10.1 4.562 4.570 4.562 4.570 dbcsr_mm_accdrv_process 8119 10.0 0.439 0.493 4.173 4.191 dbcsr_copy_into_existing 5 8.0 3.886 3.920 3.886 3.920 dbcsr_mm_accdrv_process_sort 8119 11.0 3.674 3.685 3.674 3.685 mp_waitall_1 2666 10.6 3.506 3.625 3.506 3.625 dbcsr_mm_multrec_init 95 8.2 0.000 0.000 3.518 3.605 dbcsr_mm_csr_init 95 9.2 0.006 0.006 3.518 3.605 dbcsr_mm_sched_init 95 10.2 0.000 0.000 3.489 3.576 dbcsr_mm_accdrv_init 95 11.2 0.245 0.290 3.489 3.575 tree_to_linear_d 11 10.5 3.552 3.565 3.552 3.565 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="bench_dftb", label="bench_dftb", y=162.373, yerr=0.0 Plot: name="bench_dftb_timings_6cpu_1gpu", title="Timings of bench_dftb with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="rest", label="rest", y=81.53899999999999, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="multiply_cannon_multrec", label="multiply_cannon_multrec", y=43.368, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="dbcsr_complete_redistribute", label="dbcsr_complete_redistribute", y=10.776, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="dbcsr_matrix_vector_mult_local", label="dbcsr_matrix_vector_mult_local", y=9.902, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="tb_ewald_overlap", label="tb_ewald_overlap", y=8.988, yerr=0.0 PlotPoint: plot="bench_dftb_timings_6cpu_1gpu", name="calculate_norms", label="calculate_norms", y=7.8, yerr=0.0 Running dbcsr.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/dbcsr_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.005 0.006 49.648 49.648 lib_test 1 2.0 0.000 0.000 49.634 49.641 dbcsr_run_tests 3 3.0 0.000 0.000 49.634 49.640 test_multiplies_multiproc 3 4.0 0.001 0.001 38.527 38.539 dbcsr_multiply_generic 9 5.0 0.002 0.002 29.850 29.851 multiply_cannon 9 6.0 0.269 0.348 19.668 20.216 multiply_cannon_loop 9 7.0 0.003 0.003 18.253 18.667 multiply_cannon_multrec 18 8.0 9.806 10.294 16.886 17.292 dbcsr_make_random_matrix 9 4.0 7.672 7.682 10.969 10.984 dbcsr_finalize 27 5.7 0.001 0.001 7.503 7.521 dbcsr_merge_all 18 6.5 3.720 3.725 7.386 7.407 dbcsr_mm_accdrv_process 8199 9.0 1.134 1.157 6.861 6.943 dbcsr_redistribute 9 5.0 3.616 3.648 6.053 6.063 make_m2s 18 6.0 0.001 0.001 5.157 5.158 make_images 18 7.0 0.366 0.366 5.122 5.122 dbcsr_mm_accdrv_process_sort 8199 10.0 4.720 4.725 4.720 4.725 make_images_data 18 8.0 0.001 0.001 3.049 3.056 hybrid_alltoall_any 18 9.0 2.520 2.524 3.008 3.015 mp_alltoall_d11v 27 6.0 2.145 2.148 2.145 2.148 tree_to_linear_d 9 7.0 1.872 1.889 1.872 1.889 dbcsr_data_copy_aa2 18 7.5 1.653 1.661 1.653 1.661 dbcsr_data_release 507 7.7 1.431 1.449 1.431 1.449 mp_sum_l 61 4.9 0.563 1.120 0.563 1.120 dbcsr_multiply_generic_mpsum_f 9 6.0 0.000 0.000 0.563 1.119 dbcsr_data_new 354 7.4 0.992 1.119 0.992 1.119 jit_kernel_multiply 6 10.0 1.007 1.118 1.007 1.118 dbcsr_checksum 6 5.0 1.055 1.074 1.074 1.074 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="dbcsr", label="dbcsr", y=49.648, yerr=0.0 Plot: name="dbcsr_timings_6cpu_1gpu", title="Timings of dbcsr with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="rest", label="rest", y=20.114000000000004, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="multiply_cannon_multrec", label="multiply_cannon_multrec", y=9.806, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_make_random_matrix", label="dbcsr_make_random_matrix", y=7.672, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_mm_accdrv_process_sort", label="dbcsr_mm_accdrv_process_sort", y=4.72, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_merge_all", label="dbcsr_merge_all", y=3.72, yerr=0.0 PlotPoint: plot="dbcsr_timings_6cpu_1gpu", name="dbcsr_redistribute", label="dbcsr_redistribute", y=3.616, yerr=0.0 Running MQAE_single_node.inp with 3 threads and 2 ranks... done. From /workspace/artifacts/MQAE_single_node_6cpu_1gpu.out: ------------------------------------------------------------------------------- - - - T I M I N G - - - ------------------------------------------------------------------------------- SUBROUTINE CALLS ASD SELF TIME TOTAL TIME MAXIMUM AVERAGE MAXIMUM AVERAGE MAXIMUM CP2K 1 1.0 0.041 0.042 210.282 210.282 qs_mol_dyn_low 1 2.0 0.004 0.004 208.638 208.676 qs_forces 6 3.8 0.001 0.001 132.366 132.366 qs_energies 6 4.8 0.001 0.001 124.906 124.906 scf_env_do_scf 6 5.8 0.000 0.001 117.887 117.887 scf_env_do_scf_inner_loop 113 6.2 0.006 0.008 110.475 110.475 velocity_verlet 5 3.0 0.003 0.003 99.468 99.518 rebuild_ks_matrix 119 8.1 0.001 0.001 91.032 91.033 qs_ks_build_kohn_sham_matrix 119 9.1 0.026 0.026 91.032 91.032 qs_ks_update_qs_env 119 7.3 0.001 0.001 85.933 85.934 fft_wrap_pw1pw2 2059 12.4 0.045 0.045 72.124 72.141 fft_wrap_pw1pw2_150 1321 13.9 0.009 0.009 69.186 69.257 qs_vxc_create 119 10.1 0.002 0.002 57.638 57.639 xc_vxc_pw_create 119 11.1 1.572 1.577 57.636 57.637 xc_pw_derive 714 13.1 0.010 0.010 40.357 40.366 qmmm_el_coupling 6 3.8 0.000 0.000 40.117 40.118 qmmm_elec_with_gaussian 6 4.8 0.036 0.037 40.111 40.113 qmmm_elec_with_gaussian_low 6 5.8 0.000 0.000 38.225 38.662 pw_gpu_c1dr3d_3d_ps 1095 14.8 10.728 10.823 38.600 38.609 qmmm_elec_gaussian_low_G 6 6.8 33.370 33.786 33.370 33.786 pw_gpu_r3dc1d_3d_ps 964 14.0 9.773 9.847 33.467 33.494 qmmm_forces 6 3.8 0.001 0.001 33.353 33.353 qmmm_forces_with_gaussian 6 4.8 0.049 0.050 32.673 32.893 qmmm_force_with_gaussian_low 6 5.8 0.000 0.000 31.202 31.416 xc_rho_set_and_dset_create 119 12.1 2.473 2.476 28.604 28.623 xc_pw_divergence 119 12.1 0.005 0.006 27.051 27.058 qmmm_forces_gaussian_low_G 6 6.8 26.087 26.359 26.087 26.359 qs_rho_update_rho_low 119 7.3 0.001 0.001 23.964 24.193 calculate_rho_elec 119 8.3 1.129 1.131 23.963 24.193 mp_alltoall_z22v 2059 16.4 18.331 18.424 18.331 18.424 density_rs2pw 119 9.3 0.008 0.008 17.678 17.920 sum_up_and_integrate 119 10.1 0.005 0.005 16.635 16.651 integrate_v_rspace 119 11.1 0.022 0.022 16.446 16.462 x_to_yz 1095 15.8 2.340 2.350 12.352 12.370 dbcsr_multiply_generic 2616 12.3 0.107 0.110 11.565 11.728 potential_pw2rs 119 12.1 0.034 0.035 10.701 10.704 yz_to_x 964 15.0 1.798 1.814 10.117 10.187 multiply_cannon 2616 13.3 0.242 0.243 9.796 10.061 multiply_cannon_loop 2616 14.3 0.281 0.284 9.268 9.530 qs_ks_ddapc 119 10.1 0.002 0.003 9.482 9.508 pw_gpu_sf 1095 15.8 8.804 8.840 8.804 8.840 pw_gpu_fg 964 15.0 8.516 8.591 8.516 8.591 init_scf_loop 6 6.8 0.000 0.000 7.409 7.409 qs_scf_new_mos 113 7.2 0.001 0.001 7.400 7.401 qs_scf_loop_do_ot 113 8.2 0.001 0.001 7.399 7.400 ot_scf_mini 113 9.2 0.002 0.002 7.100 7.102 multiply_cannon_multrec 5232 15.3 3.242 3.282 6.807 6.845 pw_gpu_ffc 1095 15.8 6.698 6.730 6.698 6.730 grid_integrate_task_list 119 12.1 5.723 5.740 5.723 5.740 xc_functional_eval 238 13.1 0.003 0.003 5.369 5.380 qmmm_forces_gaussian_low_R 6 6.8 0.000 0.000 5.116 5.174 qmmm_forces_with_gaussian_LG 6 7.8 5.116 5.174 5.116 5.174 qs_ks_update_qs_env_forces 6 4.8 0.000 0.000 5.132 5.132 grid_collocate_task_list 119 9.3 5.121 5.123 5.121 5.123 pw_gpu_cff 964 15.0 4.993 5.037 4.993 5.037 ot_mini 113 10.2 0.001 0.001 4.900 4.900 qmmm_elec_gaussian_low_R 6 6.8 0.000 0.000 4.855 4.876 qmmm_elec_with_gaussian_LG 6 7.8 4.855 4.876 4.855 4.876 pw_poisson_solve 125 9.9 0.003 0.004 4.819 4.823 init_scf_run 6 5.8 0.000 0.000 4.620 4.620 scf_env_initial_rho_setup 6 6.8 0.000 0.000 4.620 4.620 ------------------------------------------------------------------------------- PlotPoint: plot="total_timings_6cpu_1gpu", name="MQAE_single_node", label="MQAE_single_node", y=210.282, yerr=0.0 Plot: name="MQAE_single_node_timings_6cpu_1gpu", title="Timings of MQAE_single_node with 6 CPU Cores and 1 GPU", ylabel="time [s]" PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="rest", label="rest", y=111.99300000000001, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="qmmm_elec_gaussian_low_G", label="qmmm_elec_gaussian_low_G", y=33.37, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="qmmm_forces_gaussian_low_G", label="qmmm_forces_gaussian_low_G", y=26.087, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="mp_alltoall_z22v", label="mp_alltoall_z22v", y=18.331, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="pw_gpu_c1dr3d_3d_ps", label="pw_gpu_c1dr3d_3d_ps", y=10.728, yerr=0.0 PlotPoint: plot="MQAE_single_node_timings_6cpu_1gpu", name="pw_gpu_r3dc1d_3d_ps", label="pw_gpu_r3dc1d_3d_ps", y=9.773, yerr=0.0 Summary: Performance test took 42 minutes. Status: OK ---> Removed intermediate container e30946923835 ---> 8d75ef35cdb0 Step 46/47 : CMD cat $(find ./report.log -mmin +10) | sed '/^Summary:/ s/$/ (cached)/' ---> Running in d730b3ecd358 ---> Removed intermediate container d730b3ecd358 ---> 345f5e43cfe6 Step 47/47 : ENTRYPOINT [] ---> Running in 2df6c308ce91 ---> Removed intermediate container 2df6c308ce91 ---> d5ce4c36b277 [Warning] One or more build-args [GIT_COMMIT_SHA SPACK_CACHE] were not consumed Successfully built d5ce4c36b277 Successfully tagged us-central1-docker.pkg.dev/cp2k-org-project/cp2kci/img_cp2k-perf-cuda-volta:master Pushing new image... done. #################### Running Image cp2k-perf-cuda-volta #################### Uploading artifacts... done EndDate: 2026-10-06 02:11:44+00:00