AMSS-NCKU

64-BitBrainstorm_2026/AMSS-NCKU

Author	SHA1	Message	Date
CGH0S7	6fffaa13f6	Optimize buffer_width dynamically based on FD order to improve scalability	2026-01-31 19:04:19 +08:00
CGH0S7	6684016e8c	Optimize MPI domain decomposition min_width calculation to improve scalability	2026-01-31 16:23:16 +08:00
CGH0S7	d11eaa2242	Optimize bssn_rhs.f90: Fuse loops for metric inversion and Christoffel symbols to improve cache locality	2026-01-21 11:22:33 +08:00
CGH0S7	ef96766e22	优化 compute_rhs_bssn 热点路径并加入 NaN 检查开关 - 用 DEBUG_NAN_CHECK 宏按需启用 NaN 检查，并在输入/宏生成器中新增 Debug_NaN_Check 配置 - 逆度量改为先求行列式再乘法展开，减少除法；并在 Gam^i/Christoffel 处提取公共子表达式 - 预置批量 fderivs 辅助例程，便于后续矢量化/合并导数计算 - 将默认 MPI_processes 调整为 8 变更涉及： - AMSS_NCKU_source/bssn_rhs.f90 - generate_macrodef.py - AMSS_NCKU_Input.py - AMSS_NCKU_Input_Mini.py - inputfile_example/AMSS_NCKU_Input.py - AMSS_NCKU_source/diff_new.f90 TODO: fmisc.f90 polint()	2026-01-20 19:37:26 +08:00
CGH0S7	ae7b77e44c	Setup GW150914-mini test case for laptop development - Add AMSS_NCKU_Input_Mini.py with reduced grid resolution and MPI processes - Add AMSS_NCKU_MiniProgram.py launcher with automatic configuration swapping - Update makefile_and_run.py to reduce build jobs and remove CPU binding for laptop - Update .gitignore to exclude GW150914-mini output directory	2026-01-20 00:31:40 +08:00
CGH0S7	26c81d8e81	makefile updated	2026-01-19 23:53:16 +08:00
CGH0S7	9deeda9831	Refactor verification method and optimize numerical kernels with oneMKL BLAS This commit transitions the verification approach from post-Newtonian theory comparison to regression testing against baseline simulations, and optimizes critical numerical kernels using Intel oneMKL BLAS routines. Verification Changes: - Replace PN theory-based RMS calculation with trajectory-based comparison - Compare optimized results against baseline (GW150914-origin) on XY plane - Compute RMS independently for BH1 and BH2, report maximum as final metric - Update documentation to reflect new regression test methodology Performance Optimizations: - Replace manual vector operations with oneMKL BLAS routines: * norm2() and scalarproduct() now use cblas_dnrm2/cblas_ddot (C++) * L2 norm calculations use DDOT for dot products (Fortran) * Interpolation weighted sums use DDOT (Fortran) - Disable OpenMP threading (switch to sequential MKL) for better performance Build Configuration: - Switch from lmkl_intel_thread to lmkl_sequential - Remove -qopenmp flags from compiler options - Maintain aggressive optimization flags (-O3, -xHost, -fp-model fast=2, -fma) Other Changes: - Update .gitignore for GW150914-origin, docs, and temporary files	2026-01-18 14:25:21 +08:00
CGH0S7	3a7bce3af2	Update Intel oneAPI configuration and CPU binding settings - Update makefile.inc with Intel oneAPI compiler flags and oneMKL linking - Configure taskset CPU binding to use nohz_full cores (4-55, 60-111) - Set build parallelism to 104 jobs for faster compilation - Update MPI process count to 48 in input configuration	2026-01-17 20:41:02 +08:00
CGH0S7	c6945bb095	Rename verify_accuracy.py to AMSS_NCKU_Verify_ASC26.py and improve visual output	2026-01-17 14:54:33 +08:00
CGH0S7	0d24f1503c	Add accuracy verification script for GW150914 simulation - Verify RMS error < 1% (black hole trajectory vs. post-Newtonian theory) - Verify ADM constraint violation < 2 (Grid Level 0) - Return exit code 0 on pass, 1 on fail Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-17 00:37:30 +08:00
CGH0S7	cb252f5ea2	Optimize numerical algorithms with Intel oneMKL - FFT.f90: Replace hand-written Cooley-Tukey FFT with oneMKL DFTI - ilucg.f90: Replace manual dot product loop with BLAS DDOT - gaussj.C: Replace Gauss-Jordan elimination with LAPACK dgesv/dgetri - makefile.inc: Add MKL include paths and library linking All optimizations maintain mathematical equivalence and numerical precision.	2026-01-16 10:58:11 +08:00
CGH0S7	7a76cbaafd	Add numactl CPU binding to avoid cores 0-3 and 56-59 Bind all computation processes (ABE, ABEGPU, TwoPunctureABE) to CPU cores 4-55 and 60-111 using numactl --physcpubind to prevent interference with system processes on reserved cores.	2026-01-16 10:24:46 +08:00
CGH0S7	57a7376044	Switch compiler toolchain from GCC to Intel oneAPI - makefile.inc: Replace GCC compilers with Intel oneAPI - C/C++: gcc/g++ -> icx/icpx - Fortran: gfortran -> ifx - MPI linker: mpic++ -> mpiicpx - Update LDLIBS and compiler flags accordingly - macrodef.h: Fix include path (microdef.fh -> macrodef.fh) Requires: source /home/intel/oneapi/setvars.sh before build	2026-01-15 16:32:12 +08:00
CGH0S7	cd5ceaa15f	main branch updated	2026-01-14 08:55:53 +08:00
CGH0S7	f2fc9af70e	asc26 amss-ncku initialized	2026-01-13 15:01:15 +08:00

15 Commits