Skip to content

Performance Report

Time-grid benchmark

python benchmarks/benchmark_engine.py measures the complete engine at 1 and 1,000 cells on the default 70-year support and a four-point explicit grid. The current support contains 289 nodes for the June 1, 2001 event; the custom grid contains 4. Both runs return 2 cohorts, 54 materials, 19 plants, and 11 animal products.

Cells Grid Nodes Wall time Cells/second Peak traced memory Process RSS high-water
1 VBA default, 70 years 289 1.411 s 0.7 13.1 MB 85.1 MB
1 Custom, 4 points 4 0.548 s 1.8 1.5 MB 85.1 MB
1,000 VBA default, 70 years 289 9.291 s 107.6 1,650.8 MB 1,816.1 MB
1,000 Custom, 4 points 4 0.695 s 1,438.2 27.0 MB 1,816.1 MB

RSS is the process-level VmHWM high-water mark in decimal MB; the cases run sequentially in one benchmark process, so each value includes earlier cases. Traced memory is measured separately for each run.

Chunked and unchunked execution were compared at 1,000 cells using 100-cell chunks for both support grids. Per-capita interval dose and material concentrations compare within rtol=1e-13 (zero absolute tolerance). The standalone performance test also compares a 70-year, three-cell run with 2-/1-cell chunks. Wall-time values are observations, not acceptance thresholds.

The last 100,000-cell benchmark remains a historical three-point run: 49.043 s and about 2,735 MB traced peak. It is not evidence that the 100,000-cell 70-year case is safe to allocate.

Findings

The current run evaluates hay preparation and Type-5 crop harvest averaging on fixed daily source grids and evaluates plant-derived storage edges at exact shifted source times. These correctness changes increase work and peak memory relative to the earlier reporting-grid interpolation implementation. The explicit iter_chunks API bounds leading-batch memory; run() and run_default() still construct a full in-memory result.

F-01: Phenology Temporary Arrays

  • Function: ecosys.kernels.plants.evaluate_phenology
  • Evidence: LAI and biomass are assembled with separate tuple comprehensions and np.stack, producing full-size intermediate arrays at the 100,000-cell benchmark size.
  • Intended change: evaluate into preallocated final-axis arrays or fuse the curve assembly if profiling shows memory pressure in full engine runs.
  • Required exactness tests: tests/kernels/test_phenology.py, scalar/raster equivalence, leap-day behavior, and the 100,000-cell benchmark.
  • Resolution: P11-05B-01 preallocated the final LAI/biomass arrays and reduced kernel peak traced memory while preserving numerical tests.

F-02: Full-Engine Peak Memory At 100,000 Cells

  • Function: EcosysEngine.run()
  • Evidence: the complete tree materializes batch + time + entity arrays for plants (19), materials (54), and animal products (11), plus dose and pathway intermediates. The latest daily-harvest correctness path reaches about 2.74 GB of traced peak memory at 100,000 cells.
  • Intended change: if a deployment requires lower memory, process leading batch axes in chunks and concatenate; chunked and unchunked values already agree exactly.
  • Resolution: EcosysEngine.iter_chunks streams leading-batch results and the 100,000-cell, 70-year workload was verified through that bounded API. run() does not automatically chunk its full result. No per-cell production loop exists; chunk equivalence is covered by the benchmark and performance tests.

No per-cell Python loops were found in the benchmarked production paths. The existing loops iterate over fixed parameter entities or time intervals and are not candidates for batch-loop removal.

Bounded streaming at 100,000 cells

At 1,000 cells and 289 support nodes, one batch + time + entity array uses about 124.8 MB for 54 materials, 44.0 MB for 19 plants, or 25.4 MB for 11 animal products. The complete result retains both per-event and superposed concentrations, about 388 MB for those concentration outputs alone before dose and temporary arrays. The measured 1,650.8 MB traced peak suggested an unbounded 100,000-cell run could exceed 100 GB.

EcosysEngine.iter_chunks(request, chunk_size=...) now streams each flattened leading-batch slice as an independent SimulationResult; the consumer can persist or reduce each chunk and release it before the next. The complete 100,000-cell, 70-year support was streamed in 40 chunks of 2,500 cells:

Cells Nodes Chunk size Wall time Peak traced memory Process peak RSS
100,000 289 2,500 886.513 s 3.825 GB 3.933 GB

Reproduce with python benchmarks/benchmark_engine.py --stream-100000 --stream-chunk-size 2500. The benchmark consumes each chunk, checks ordered coverage of the full batch, and reports a dose checksum of 0.0095453 Sv without concatenating the result tree. run() and run_default() still return complete in-memory trees; large raster callers should use iter_chunks and consume results incrementally. The 100,000-cell streaming workload is verified, while direct full-result allocation at that size is not advertised.