Performance Report
Time-grid benchmark
python benchmarks/benchmark_engine.py measures the complete engine at 1 and
1,000 cells on the default 70-year support and a four-point explicit grid. The
current support contains 289 nodes for the June 1, 2001 event; the custom grid
contains 4. Both runs return 2 cohorts, 54 materials, 19 plants, and 11 animal
products.
| Cells | Grid | Nodes | Wall time | Cells/second | Peak traced memory | Process RSS high-water |
|---|---|---|---|---|---|---|
| 1 | VBA default, 70 years | 289 | 1.411 s | 0.7 | 13.1 MB | 85.1 MB |
| 1 | Custom, 4 points | 4 | 0.548 s | 1.8 | 1.5 MB | 85.1 MB |
| 1,000 | VBA default, 70 years | 289 | 9.291 s | 107.6 | 1,650.8 MB | 1,816.1 MB |
| 1,000 | Custom, 4 points | 4 | 0.695 s | 1,438.2 | 27.0 MB | 1,816.1 MB |
RSS is the process-level VmHWM high-water mark in decimal MB; the cases run
sequentially in one benchmark process, so each value includes earlier cases.
Traced memory is measured separately for each run.
Chunked and unchunked execution were compared at 1,000 cells using 100-cell
chunks for both support grids. Per-capita interval dose and material
concentrations compare within rtol=1e-13 (zero absolute tolerance). The
standalone performance test also compares a 70-year, three-cell run with
2-/1-cell chunks. Wall-time values are observations, not acceptance thresholds.
The last 100,000-cell benchmark remains a historical three-point run: 49.043 s and about 2,735 MB traced peak. It is not evidence that the 100,000-cell 70-year case is safe to allocate.
Findings
The current run evaluates hay preparation and Type-5 crop harvest averaging on
fixed daily source grids and evaluates plant-derived storage edges at exact
shifted source times. These correctness changes increase work and peak memory
relative to the earlier reporting-grid interpolation implementation. The
explicit iter_chunks API bounds leading-batch memory; run() and
run_default() still construct a full in-memory result.
F-01: Phenology Temporary Arrays
- Function:
ecosys.kernels.plants.evaluate_phenology - Evidence: LAI and biomass are assembled with separate tuple comprehensions
and
np.stack, producing full-size intermediate arrays at the 100,000-cell benchmark size. - Intended change: evaluate into preallocated final-axis arrays or fuse the curve assembly if profiling shows memory pressure in full engine runs.
- Required exactness tests:
tests/kernels/test_phenology.py, scalar/raster equivalence, leap-day behavior, and the 100,000-cell benchmark. - Resolution:
P11-05B-01preallocated the final LAI/biomass arrays and reduced kernel peak traced memory while preserving numerical tests.
F-02: Full-Engine Peak Memory At 100,000 Cells
- Function:
EcosysEngine.run() - Evidence: the complete tree materializes
batch + time + entityarrays for plants (19), materials (54), and animal products (11), plus dose and pathway intermediates. The latest daily-harvest correctness path reaches about 2.74 GB of traced peak memory at 100,000 cells. - Intended change: if a deployment requires lower memory, process leading batch axes in chunks and concatenate; chunked and unchunked values already agree exactly.
- Resolution:
EcosysEngine.iter_chunksstreams leading-batch results and the 100,000-cell, 70-year workload was verified through that bounded API.run()does not automatically chunk its full result. No per-cell production loop exists; chunk equivalence is covered by the benchmark and performance tests.
No per-cell Python loops were found in the benchmarked production paths. The existing loops iterate over fixed parameter entities or time intervals and are not candidates for batch-loop removal.
Bounded streaming at 100,000 cells
At 1,000 cells and 289 support nodes, one batch + time + entity array uses
about 124.8 MB for 54 materials, 44.0 MB for 19 plants, or 25.4 MB for 11 animal
products. The complete result retains both per-event and superposed
concentrations, about 388 MB for those concentration outputs alone before dose
and temporary arrays. The measured 1,650.8 MB traced peak suggested an
unbounded 100,000-cell run could exceed 100 GB.
EcosysEngine.iter_chunks(request, chunk_size=...) now streams each flattened
leading-batch slice as an independent SimulationResult; the consumer can
persist or reduce each chunk and release it before the next. The complete
100,000-cell, 70-year support was streamed in 40 chunks of 2,500 cells:
| Cells | Nodes | Chunk size | Wall time | Peak traced memory | Process peak RSS |
|---|---|---|---|---|---|
| 100,000 | 289 | 2,500 | 886.513 s | 3.825 GB | 3.933 GB |
Reproduce with python benchmarks/benchmark_engine.py --stream-100000
--stream-chunk-size 2500. The benchmark consumes each chunk, checks ordered
coverage of the full batch, and reports a dose checksum of 0.0095453 Sv without
concatenating the result tree. run() and run_default() still return complete in-memory
trees; large raster callers should use iter_chunks and consume results
incrementally. The 100,000-cell streaming workload is verified, while direct
full-result allocation at that size is not advertised.