Domain benchmark dossier

SpatialAgents Benchmark 08

Integrated multi-hazard and climate-risk analysis for Stuttgart—planned, calculated, visualised and documented by 7 agent-model systems with SpatialAgents and QGIS, including one local model on a notebook.

QGIS project from SpatialAgents Benchmark 08 with hazard layers, exposure analysis and evacuation routes
Benchmark 08 / QGISThe generated project connects 30 m terrain, six hazard indices, three geometry types, climate scenarios and routes.
  • 7model series
  • 21assessed runs
  • 87/100reference run

Model comparison · Benchmark 08

21 runs in direct comparison.

Individual values show variation across 7 agent-model systems; mean, median, range and population standard deviation use every approved run.

Individual run scores, range and mean for each agent-model system on a shared scale from 0 to 100 points.
Model comparison with individual runsIndividual run scores, range and mean for each agent-model system on a shared scale from 0 to 100 points.

Exact run scores and statistics

ModelRun valuesMeanMedianMinimumMaximumStandard deviation
Claude Opus 5Claude Code · Cloud92 · 91 · 9091.091.090920.8
Claude Sonnet 5Claude Code · Cloud85 · 85 · 8685.385.085860.5
Qwen3.8-Flash-NextOpenCode · local · HP ZBook81 · 83 · 8783.783.081872.5
GPT-5.6 SolCodex CLI · Cloud78 · 83 · 8381.383.078832.4
GPT-5.6 TerraCodex CLI · Cloud78 · 75 · 7977.378.075791.7
GPT-5.6 LunaCodex CLI · Cloud71 · 72 · 7572.772.071751.7
Claude Haiku 4.5Claude Code · Cloud53 · 59 · 5254.753.052593.1
Mean contributions of the six assessment areas for each agent-model system.
Score compositionMean contributions of the six assessment areas for each agent-model system.

Six assessment areas per run

Values are weighted contributions before final rounding of the total score.

ModelRunScoreHard checks passedHard artefact and completeness checksRequired SpatialAgents skillsDomain rubricVisual agreementTargeted tool selectionAvoidance of redundant calls
Claude Opus 519215/1525.010.035.018.23.30.0
Claude Opus 529115/1525.010.035.016.35.00.0
Claude Opus 539015/1525.010.033.517.73.30.0
Claude Sonnet 528515/1525.010.031.515.03.30.0
Claude Sonnet 538515/1525.010.032.513.73.30.0
Claude Sonnet 548615/1525.010.031.014.75.00.0
Qwen3.8-Flash-Next18115/1525.010.029.511.85.00.0
Qwen3.8-Flash-Next28315/1525.010.031.013.63.30.0
Qwen3.8-Flash-Next38715/1525.010.032.514.65.00.0
GPT-5.6 Sol17815/1525.010.033.09.70.00.0
GPT-5.6 Sol28315/1525.010.033.015.10.00.0
GPT-5.6 Sol38315/1525.010.032.015.60.00.0
GPT-5.6 Terra17815/1525.010.032.010.80.00.0
GPT-5.6 Terra27515/1525.010.031.09.20.00.0
GPT-5.6 Terra37915/1525.010.031.512.80.00.0
GPT-5.6 Luna17115/1525.010.029.56.30.00.0
GPT-5.6 Luna27215/1525.010.033.03.80.00.0
GPT-5.6 Luna37515/1525.010.030.59.20.00.0
Claude Haiku 4.515313/1521.710.015.00.15.01.0
Claude Haiku 4.525914/1523.310.016.50.25.04.0
Claude Haiku 4.535212/1520.010.011.50.65.05.0
Mean completion of all eleven domain criteria across every approved run.
Domain criteria comparedMean completion of all eleven domain criteria across every approved run.

Criterion means with run values

Each cell shows the mean first and every approved raw run score underneath.

CriterionMax.Claude Opus 5Claude Sonnet 5Qwen3.8-Flash-NextGPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaClaude Haiku 4.5
Initial plan present and traceable55.00runs 5 · 5 · 54.67runs 4 · 5 · 55.00runs 5 · 5 · 55.00runs 5 · 5 · 55.00runs 5 · 5 · 54.33runs 4 · 4 · 53.33runs 4 · 2 · 4
Geospatial API used consistently1010.00runs 10 · 10 · 109.67runs 10 · 9 · 107.67runs 5 · 9 · 910.00runs 10 · 10 · 109.67runs 10 · 10 · 910.00runs 10 · 10 · 102.33runs 3 · 3 · 1
Six physically plausible hazard indices1010.00runs 10 · 10 · 109.33runs 9 · 9 · 108.67runs 8 · 9 · 99.00runs 9 · 9 · 99.00runs 9 · 9 · 98.33runs 8 · 9 · 85.33runs 5 · 7 · 4
Clean multi-geometry sampling88.00runs 8 · 8 · 88.00runs 8 · 8 · 88.00runs 8 · 8 · 88.00runs 8 · 8 · 87.67runs 8 · 7 · 87.33runs 8 · 8 · 63.33runs 2 · 5 · 3
HAND inundation and statistics55.00runs 5 · 5 · 54.33runs 4 · 4 · 54.33runs 5 · 4 · 44.33runs 5 · 5 · 35.00runs 5 · 5 · 54.00runs 3 · 5 · 41.00runs 1 · 2 · 0
Climate projection and class shift76.67runs 7 · 7 · 66.67runs 7 · 7 · 66.67runs 7 · 7 · 67.00runs 7 · 7 · 76.33runs 6 · 7 · 67.00runs 7 · 7 · 74.33runs 5 · 5 · 3
Evacuation routes with vulnerability analysis55.00runs 5 · 5 · 54.33runs 5 · 4 · 44.67runs 5 · 4 · 54.67runs 5 · 5 · 44.33runs 5 · 4 · 44.00runs 3 · 4 · 51.00runs 1 · 1 · 1
Charts communicate findings54.67runs 5 · 5 · 44.33runs 4 · 5 · 44.33runs 4 · 4 · 53.67runs 3 · 4 · 43.33runs 3 · 3 · 43.67runs 4 · 4 · 33.00runs 3 · 3 · 3
Organised QGIS layer tree54.67runs 5 · 5 · 44.00runs 4 · 5 · 33.33runs 3 · 2 · 54.00runs 4 · 4 · 42.67runs 3 · 2 · 33.33runs 2 · 5 · 32.00runs 3 · 2 · 1
Complete results report55.00runs 5 · 5 · 54.67runs 5 · 5 · 45.00runs 5 · 5 · 55.00runs 5 · 5 · 55.00runs 5 · 5 · 55.00runs 5 · 5 · 53.00runs 3 · 3 · 3
Documented visual self-review55.00runs 5 · 5 · 53.33runs 3 · 4 · 34.33runs 4 · 5 · 44.67runs 5 · 4 · 55.00runs 5 · 5 · 55.00runs 5 · 5 · 50.00runs 0 · 0 · 0

Time and tokens

It is not the geoprocessing that costs time — it is the writing.

The time split is reconstructed from the adapter traces: model time, geoprocessing, QGIS control and file access have to add up to the wall clock.

Mean runtime per model, split into model time, geoprocessing and remaining tool time; writing speed in tokens per second on the right.
Where the runtime goesMean runtime per model, split into model time, geoprocessing and remaining tool time; writing speed in tokens per second on the right.

On average the local model spends 8.24 minutes on geoprocessing, ahead of Opus 5. Its wall-clock gap comes from writing: 32 tokens per second including prefill (43 in pure generation) against 86 to 103 for the Claude models. QGIS rendering costs every measured model less than half a minute in total.

Runtime per run

Claude Haiku 4.5 needs 9.4 minutes per run on average, Qwen3.8-Flash-Next 2 hours and 11 minutes. The other cloud models lie between 33 and 56 minutes. The time balance above splits the runtime for four systems; this table gives the wall clock for all seven.

ModelRuntime per run in minutes (min–max)MeanRuns
Claude Opus 534.0–52.841.33
Claude Sonnet 553.8–58.656.33
Qwen3.8-Flash-Next85.7–162.8130.53
GPT-5.6 Sol29.2–61.642.03
GPT-5.6 Terra32.2–35.933.83
GPT-5.6 Luna28.2–36.333.43
Claude Haiku 4.58.0–10.39.43

Wall-clock time per run from the score export, read on 19 September 2026; for all seven systems, without the split into model and geoprocessing time.

Tool calls per run

The range runs from 31 to 471 calls per run. The longest run needed 2 hours and 43 minutes for 341 calls. The runs were non-interactive: asking a human was technically impossible. Permission to work is granted once at the start, not at every step.

ModelTool calls per run (min–max)MeanRuns
Claude Opus 5280–288284.03
Claude Sonnet 5307–403358.73
Qwen3.8-Flash-Next341–471392.03
GPT-5.6 Sol209–264232.73
GPT-5.6 Terra201–249218.73
GPT-5.6 Luna177–256223.33
Claude Haiku 4.531–9260.73

Tool calls per run from the score export, counted on 18 September 2026; the runs were non-interactive, so a question back to the user was technically impossible.

Mean token consumption per run in millions for each agent-model system.
Tokens per runMean token consumption per run in millions for each agent-model system.

More tokens do not buy better maps: Sonnet 5 consumes five times as much as the GPT-5.6 variants and lands four points above them; the local model's consumption is roughly level with Opus 5, and it was the only one that had to compact its context.

Cost balance

What a run costs.

Cost per run is calculated from the token counts — Anthropic rates from the vendor documentation, GPT-5.6 rates from two cross-checked rate overviews; the local run pays only for electricity.

Mean cost per run in US dollars for each agent-model system, split by cost item; the local run pays only for electricity.
What a run costsMean cost per run in US dollars for each agent-model system, split by cost item; the local run pays only for electricity.

You are not paying for the writing but for the re-reading: for Sonnet 5, 81 percent of the run cost goes to reading its own context and 9 percent to text output. Per point, Opus 5 and Sonnet 5 cost the same — 0.28 against 0.27 dollars — while Opus scores 5.7 points higher. The local run costs 0.06 dollars of electricity; the roughly 4,000-dollar notebook is not included in that figure and equals about 160 Opus runs.

Without discounts and without the cost of the assessment itself. The GPT-5.6 runs of 19 August 2026 are calculated at rates of 18 September 2026. Electricity is set at 0.35 EUR per kilowatt hour; the 70 watts come from the machine comparison and were not measured at the power adapter.

Where the computing happens

70 watts versus 1,200.

Six of the seven models are reachable only over the network and run on accelerators their operators house in gigawatt facilities. The seventh sat on a desk during the measurement, powered by a 140-watt adapter.

Local · entirely on premises

HP ZBook Ultra G1a

Compute
Ryzen AI Max+ PRO 395 with Radeon 8060S, 70 W GPU budget, 140 W power adapter
Memory
128 GiB LPDDR5X shared between CPU and GPU — dedicated graphics memory: 0.5 GiB
Model
Qwen3.8-Flash-Next, 176.9 billion parameters, 10 of 512 experts active per token
Operation
halogen server on the local network, 262,144-token context; 886 tokens/s reading, 39.9 tokens/s writing
Provenance
own measurement on 8 September 2026 on the device that computed the runs

Data centre · Anthropic

Claude Opus 5, Sonnet 5, Haiku 4.5

Compute
AWS Trainium2 and Google TPUs in combination
Facility
Project Rainier, around 500,000 Trainium2 accelerators, reported operational by AWS on 29 October 2025
Expansion
up to 1 million Google “Ironwood” TPUs committed (Google, April 2026); one pod of this design holds 9,216 chips with 192 GB memory each
Consumption
35.6 to 94.5 million tokens per run, every request over the network
Provenance
announcements by AWS and Google; none of the parties states per-chip power draw

Data centre · OpenAI

GPT-5.6 Sol, Terra, Luna

Compute
NVIDIA accelerators, operated through Azure and the Stargate programme
Facility
Stargate: 500 billion US dollars, targeting 10 GW of grid connection (OpenAI, January 2025)
Expansion
Abilene, Texas site: over 450,000 GB200 accelerators at 1.2 GW — press figure, not confirmed by the operator
Consumption
17.7 to 19.1 million tokens per run, every request over the network
Provenance
announcement by OpenAI; per-site equipment and power draw are not published

One accelerator versus one notebook

The Blackwell B200 is the building block of the facilities above — a GB200 NVL72 rack holds 72 of them. The comparison here is one single accelerator against the whole notebook.

MetricZBook Ultra G1aNVIDIA B200Factor
Memory for the model128 GB LPDDR5X, shared between CPU and GPU186 GB HBM3e1,5 ×
Memory bandwidth256 GB/s, measured around 215 GB/s8,000 GB/s31 ×
FP16 computearound 30 TFLOPS (Radeon 8060S, 40 compute units)2,500 TFLOPS, stated without sparsity84 ×
Power draw70 W GPU budget, 140 W adapter for the whole deviceup to 1,200 W, the accelerator alone17 ×
Acquisition costaround $4,000, complete device with display and 2 TB SSD$30,000 to $55,000, a component for which vendors publish no list figures8 – 14 ×

The first row decides whether a model runs at all: a 177-billion-parameter model has to fit into memory. The notebook provides the same order of magnitude as the accelerator — it just reads it 31 times more slowly, which is where the 39.9 tokens per second come from. The task did not suffer: 87 points in the best run, one more than the best Sonnet 5 run.

B200: NVIDIA “Blackwell” datasheet (December 2024), 1,200 W variant as in the GB200 NVL72; its compute figures include sparsity, the FP16 value above is half of that. ZBook: own measurement on 8 September 2026.

Task framework

Four questions define the task.

The task connects physical hazard indices, three geometry types, statistics, climate scenarios and a structured QGIS project.

  1. 01

    Which spatial patterns do six natural hazards show across Stuttgart?

  2. 02

    How are restaurants, buildings and roads exposed to flood risk?

  3. 03

    How do the hazards correlate and how can they form a weighted composite?

  4. 04

    How does the risk picture shift under RCP 4.5 and RCP 8.5 by 2050?

Complete English translation

The assignment in full.

This editorial translation follows the complete German prompt. The German original remains authoritative and technical identifiers are unchanged.

Complete benchmark prompt

Assignment

The Office for Environmental Protection and Disaster Preparedness of the City of Stuttgart requires an integrated multi-hazard and climate-risk assessment for the urban area. Stuttgart's basin setting, the Neckar valley, the slopes at the edge of the city and the urban heat-island effect make the location particularly informative: six natural hazards have very different spatial effects here, and climate change will additionally alter the picture by 2050.

You are to carry out this analysis end to end—from data acquisition and the calculation of hazard indices to a publication-ready map composition and a written report. The assessment must answer four central questions:

  1. Topography and hazards: Which spatial patterns do the six natural hazards show across the urban area?
  2. Multi-geometry exposure: Which restaurants (points), buildings (polygons) and road sections (lines) are most exposed to flood risk?
  3. Statistics and composite assessment: How do the hazards correlate, how are their values distributed, and what does a weighted composite score look like?
  4. Climate projection: How does the risk picture shift under RCP 4.5 and RCP 8.5 by 2050?

Library recommendation and skill consultation

For geodata-based calculations, primarily use the geo-api library. Under GeospatialRaster and GeospatialVector, it brings together precisely the data sources (OSM, Microsoft Planetary Computer, DWD/CORDEX, Open-Meteo) and algorithms (terrain derivatives, spectral indices, zonal statistics, routing) required here. Read the geo-api skill before the first code action. Where geo-api does not provide a direct algorithm—for example, specialised vector operations or statistical aggregations—choose fallback packages yourself. The skills geopandas-shapely, xarray-rioxarray, rasterio, grass-gis, gdal-binaries, pyproj and qgis-styles are available and should be consulted as needed.

For QGIS visualisation (layer tree, styling, screenshots, bookmarks, layout), use the QGIS MCP bridge through the spatial-agent-bridge skill. You must also consult the workflow skill bench (Coverage Tasks section); it describes the RESULT.json schema and the screenshot self-review loop.

The choice of tools is explicitly yours—the prompt does not prescribe specific function or module names. In your initial plan, briefly explain which path you will take and why.

Procedure

Begin with a written plan of three to eight sentences in RESULT.json under summary.plan: which phases you will complete and in what order, which skills you have read and which data sources you will use. Only then begin the implementation.

Create the directory structure early:

artifacts/
├── data/
│   ├── basemaps/
│   ├── osm/                  (restaurants, buildings, roads, districts)
│   ├── terrain/              (DEM and derivatives)
│   ├── hazards/              (6 hazard indices, composite and inundation zones)
│   ├── analysis/             (sampling results and aggregations)
│   └── climate/              (RCP-scaled results)
├── charts/                   (matplotlib PNG files)
└── screenshots/              (QGIS screenshots)

Save the QGIS project as artifacts/stuttgart_hazards.qgz. Set the project CRS deliberately (EPSG:25832 is an appropriate choice for southern Germany).

Phase 1—Define the study area and acquire base data

Define the study area: the City of Stuttgart together with the surrounding ridges (Bopser, Killesberg, Frauenkopf), covering a bounding box of approximately 10 × 10 km. Obtain for this area:

  • a digital elevation model at approximately 30 m resolution (Copernicus GLO-30 via Microsoft Planetary Computer is the standard choice);
  • an OSM basemap layer;
  • OSM vector data: all restaurants (POIs), all building polygons, the drivable road network and the city-district boundaries.

DEM acquisition note: download only the area-of-interest subset, not complete tile stacks.

Organise the downloaded layers in QGIS in a meaningful layer tree, using groups such as “Terrain/Source data”, “Urban/Buildings”, “POI”, “Administration/Districts” and “Basemaps”. Style examples from the qgis-styles skill are welcome.

Create an overview screenshot at the end of the phase.

Phase 2—Terrain derivatives

The principal terrain derivatives must be produced from the DEM. At minimum, these are:

  • slope in degrees;
  • aspect from 0–360°;
  • topographic position, for example TPI at an approximately 330 m window radius;
  • topographic wetness index, for example TWI;
  • ruggedness, for example TRI;
  • profile curvature;
  • height above nearest drainage, or HAND;
  • flow accumulation;
  • hillshade for visualisation.

The library or tool used is open: geo-api provides a terrain accessor for many of these calculations, GRASS GIS offers specialised modules (see the grass-gis skill), and the GDAL CLI provides several fundamentals (see the gdal-binaries skill). HAND is usually the most demanding step, comprising flow accumulation, stream extraction and vertical distance.

Performance note: Watershed algorithms, particularly HAND and flow accumulation, can consume several tens of gigabytes of RAM for a 10 × 10 km, 30 m DEM when implemented naively in Python. If you notice that a process uses more than 20 GB RSS, stop it and change approach. GRASS GIS r.watershed is considerably more memory-efficient on large DEMs than Python alternatives because it uses time-tested C code. Alternatively, downsample the DEM to about 60 m before computing HAND; doubling the cell size quarters the pixel count. Document the chosen approach under summary.hazard_formulas.

Style the derivatives so the information is clear: hillshade as a greyscale background; slope in warm colours; TPI with a diverging valley-to-ridge scale; TWI sequentially from dry to wet; and HAND sequentially from near drainage to high above it.

Create a screenshot of the derivative overview; hillshade with a semi-transparent slope layer is a useful default.

Phase 3—Six hazard indices and a composite

Calculate six normalised hazard indices from the terrain derivatives, each scaled to [0,1], where 1 indicates the greatest hazard, and a weighted composite. The following indices and physical rationales are expected:

Index What it models Recommended components
Wind exposure Ridges and west-facing slopes are exposed to wind TPI, west-facing measure, relative elevation, slope
Frost risk Cold air pools in valleys and wet depressions Valley position (inverse TPI), TWI, low relative elevation
Flood risk Low-lying areas near drainage with large catchments Inverse HAND, TWI, flow accumulation
Heat stress Low elevation, south-facing slope and urban factor Inverse elevation, south-facing measure, urban factor (constant)
Landslide Steep, wet and concave terrain Slope, TWI, concavity
Erosion (LS factor) RUSLE LS factor from flow and slope Flow accumulation, slope function

The exact formulas and weights are at your discretion. Follow standard literature—for example RUSLE for erosion, TPI for wind and HAND for flooding—and document the selected formulas in RESULT.json under summary.hazard_formulas.

Composite: a weighted sum of the six indices. Choose and justify the weights; flood and heat would typically receive greater weight in an urban heat and flooding context.

Screenshots: create a separate map over hillshade for each signature hazard (wind, flood, heat and landslide), plus a composite overview map.

Also create HAND-based inundation zones for three scenarios: water levels of 1 m, 2 m and 5 m above the nearest drainage. Export the zones as polygons, for example to a GeoPackage. For each scenario, calculate and report inundated area in km², number of affected buildings and total length of affected roads in km.

Phase 4—Multi-geometry exposure

Three geometry types must be sampled with hazard values. This is deliberate, because each requires a different raster-to-vector extraction method:

  • Restaurants (points): extract all six hazard values plus elevation and slope for each point. Calculate a composite score for each restaurant using weights of your choice and document those weights. Assign four risk classes: Low, Medium, High and Very High.
  • Buildings (polygons): aggregate the flood index for each polygon using zonal statistics or centroid sampling, choosing the method that is appropriate. Assign four risk classes.
  • Roads (lines): similarly aggregate the flood index for each segment. Assign four risk classes and use line width proportional to the score in the visualisation.

Export all three result sets as separate GeoPackages under artifacts/data/analysis/. Also export the restaurant results as GeoJSON and CSV.

District-level aggregation: spatially join restaurants to city districts, calculate the mean, count and percentage at high risk per district, and visualise the result as choropleth polygons in QGIS.

Screenshots: one dedicated screenshot per geometry type—restaurants, buildings and roads styled by risk class—plus a district choropleth and the three inundation zones.

Phase 5—Statistics and charts

Generate at least seven charts as PNG files under artifacts/charts/. The minimum inventory is:

  1. correlation heatmap for the six hazards across all restaurants (6 × 6, Pearson, diverging colour scale with annotated values);
  2. histogram panel for the six hazards, showing the shape of each distribution;
  3. radar charts for the five restaurants with the highest risk;
  4. bar chart of restaurant count by risk class;
  5. scatter plot of elevation versus flood score, with trend line;
  6. box plots of all six hazards by risk class;
  7. climate-comparison bar chart for baseline, RCP 4.5 and RCP 8.5.

Chart 7 will be calculated in Phase 7, but its skeleton may already be prepared here.

Phase 6—High-risk inspection and evacuation routing

Select restaurants with a composite score above 0.50. Zoom to the spatial cluster with the greatest concentration and to the single restaurant with the highest score, producing a separate screenshot of each.

Select the three restaurants with the highest composite score and calculate a pedestrian route from each to a safe point, for example a point with HAND above 10 m on a ridge. Routing may use OSM: geo-api provides a routing component, or OSMnx may be used for graph-based routing; consult the corresponding skills.

Identify route segments that cross flood-prone areas with a flood index above 0.5; these are vulnerable segments. Report route length, estimated walking time and vulnerable share.

Create a screenshot of the three routes with the vulnerable segments highlighted.

Phase 7—Climate projection (RCP 4.5 and RCP 8.5, 2050 horizon)

Apply climate-scaling factors to each restaurant’s baseline hazard values. The following are indicative values for the CORDEX ensemble to 2050 in southern Germany:

Hazard RCP 4.5 factor RCP 8.5 factor Direction
Wind ~1.00 ~1.00 static
Frost ~0.77 ~0.55 decreasing
Flood ~1.03 ~1.07 slightly increasing
Heat ~1.80 ~2.59 strongly increasing
Landslide ~1.00 ~1.00 static
Erosion ~1.03 ~1.07 slightly increasing

You may use these factors directly (source: EURO-CORDEX EUR-11, MPI-M-MPI-ESM-LR, 2050) or retrieve more detailed values through geo_api.ClimateDataApi if that better suits your workflow. Document the factors used in RESULT.json.

For every restaurant, scale each hazard and cap it at 1.0, recalculate the composite and assign the new risk class. Save the results as CSV files under artifacts/data/climate/. Report the class shift as a percentage in each class for each scenario.

Generate Chart 7 from the actual data. Screenshot: display restaurants at baseline and under RCP 8.5 side by side in QGIS.

Phase 8—Finalisation

Clean up the layer tree for publication, with a clear hierarchy: analysis results at the top, hazards in the middle, and source data and basemap at the bottom.

Create a final composite screenshot with hillshade, a semi-transparent 2 m inundation zone, buildings by flood risk, restaurants by composite risk class and visible evacuation routes.

Write a report at artifacts/RESULTS.md in German or English with the usual sections: executive summary, study area, method, results for each phase, climate projections and recommendations. Embed the charts using Markdown.

Self-review loop for every screenshot

Follow skill bench, PROTOCOL §10 (Coverage Tasks): inspect every screenshot before proceeding. If you are Claude Code, open the PNG with the Read tool and inspect it visually. If you are Codex, use Bash with Python/PIL to check the pixel distribution; a standard deviation below 5 suggests an empty or uniform image. If a screenshot does not meet the expectation—because it is empty, incorrectly centred, shows the wrong visible layers or is identical to the previous phase—correct it and capture it again. Use no more than two retries per phase. Document retries in RESULT.json under summary.review_notes[].

Screenshot rules (binding)

  • After every load_project, first perform a map_navigation action such as zoom_to_layer or set_extent before taking the first screenshot. The canvas extent is undefined after loading a project in this environment and would otherwise produce an empty image.
  • Always call get_map_screenshot with explicit dimensions: width=1600, height=1000 (default 96 dpi). Window geometry is unreliable in this environment; never rely on the canvas size.
  • Never use include_overlays=true: the widget-grab path produces empty images in this environment. Map tips are not required for this task.
  • Check the content_hash returned after every screenshot. If it is identical to the preceding screenshot, the map has not changed and visibility or extent is incorrect; then apply the self-review loop.

What RESULT.json must contain

The RESULT.json must comply with the schema in the bench skill (PROTOCOL.md). At the top level it requires task_id, run_index, mode, summary, hard_checks and anti_checks. For every anti.never: rule in this task, an anti_checks[] entry with id, violation and evidence must exist. A missing entry is scored as a violation.

Under summary:

  • plan: the initial plan string described above;
  • phases_completed: array ["1", "2", ..., "8"];
  • data_sources_used: list of data-source endpoints with versions, for example {"copernicus_glo30": "dem 1.0", "osm_overpass": "<date>"};
  • hazard_formulas: dictionary with the selected formulas and weights for each hazard;
  • climate_factors_used: dictionary with scaling factors by hazard for RCP 4.5 and RCP 8.5;
  • composite_weights: dictionary with composite-score weights by geometry type;
  • screenshots_written: list of relative paths;
  • charts_written: list of relative paths;
  • inundation_stats: dictionary for each scenario, for example {"1m": {"area_km2": ..., "buildings_affected": ..., "streets_km_affected": ...}, "2m": {...}, "5m": {...}};
  • risk_class_distribution: dictionary {"baseline": {"Low": <pct>, "Medium": <pct>, "High": <pct>, "Very High": <pct>}, "rcp45": {...}, "rcp85": {...}};
  • top_10_risk_restaurants: list with name or ID, district, composite score and dominant hazard;
  • evacuation_routes: list with start restaurant, target coordinate, route length in metres, walking time in minutes and vulnerable kilometres;
  • failures: array of {phase, action, error} for each unsuccessful step;
  • review_notes: documentation of the self-review loop;
  • abort_reason: string or null.

Final cleanup

All temporary QGIS layers, groups, bookmarks and layouts belong under the tree node Benchmark/B08_stuttgart_hazard_climate_risks/run_<N> so that the next run starts from a clean state. File outputs under artifacts/ remain in place; they are the result of the task.

Assessment framework

100 points across six assessment areas.

Completeness, domain quality, visual communication and tool use are assessed separately and then combined. The runs used the institute's research system with additional domain skills for geoprocessing and the Geospatial API. How a benchmark is built and what it shows is set out in the thirteenth section of the fundamentals.

  1. 25Hard artefact and completeness checks
  2. 10Required SpatialAgents skills
  3. 35Domain rubric
  4. 20Visual agreement
  5. 5Targeted tool selection
  6. 5Avoidance of redundant calls
15 hard checks
  1. 01

    A structured result file is present.

  2. 02

    The complete result report is present.

  3. 03

    The QGIS project has been saved.

  4. 04

    At least twelve QGIS screenshots were generated.

  5. 05

    At least five charts were generated.

  6. 06

    At least three GeoPackages with analysis results are present.

  7. 07

    The inundation zones are available as a GeoPackage.

  8. 08

    The climate scenarios are available as tabular data.

  9. 09

    At least six work phases are documented.

  10. 10

    At least twelve screenshot paths are documented.

  11. 11

    At least five chart paths are documented.

  12. 12

    The hazard-index formulas are documented.

  13. 13

    Area, buildings and road length for the inundation scenarios are documented.

  14. 14

    Risk classes for the baseline and climate scenarios are documented.

  15. 15

    Three evacuation routes with metrics are documented.

Domain rubric · 70 raw points

The right-hand value is the documented Sol reference run score before weighting.

5

Initial plan present and traceable

The initial plan states the phase order, consulted skills and data sources before implementation begins.

5/5reference run
10

Geospatial API used consistently

The Geospatial API is used for acquisition and at least one analysis path; justified fallbacks remain possible.

9/10reference run
10

Six physically plausible hazard indices

All six indices are normalised, their formulas are documented and their spatial patterns are physically plausible.

9/10reference run
8

Clean multi-geometry sampling

Restaurants, buildings and roads are sampled appropriately and exported with four risk classes.

8/8reference run
5

HAND inundation and statistics

Three threshold zones and monotonic statistics for area, buildings and affected road length are present.

4/5reference run
7

Climate projection and class shift

RCP 4.5 and RCP 8.5 scores and risk classes show the expected climate-driven shift.

6/7reference run
5

Evacuation routes with vulnerability analysis

Three routes include length, walking time, vulnerable segments and a legible QGIS presentation.

5/5reference run
5

Charts communicate findings

The charts use meaningful axes, scales, titles and legends to support an interpretable statement.

5/5reference run
5

Organised QGIS layer tree

Analysis, hazards, terrain, source data and basemap form a readable hierarchy with deliberate styles.

5/5reference run
5

Complete results report

The report covers all eight phases, embeds charts, interprets findings and provides recommendations.

5/5reference run
5

Documented visual self-review

Each screenshot is inspected and any repeat capture records the correction that was made.

4/5reference run

Documented reference run

Qwen3.8-Flash-Next with OpenCode.

Run 3 forms the documented case study. Its maps, metrics and QGIS project remain bound to this run as one coherent record.

  1. 87/100score
  2. 15/15hard checks
  3. 65/70rubric points
  4. 17/17required views
  5. 7/7charts
  6. 73 %visual agreement

The SpatialAgents in the reference run used the Geospatial API for data acquisition and terrain methods, added appropriate Python and GIS tools, produced shared analysis files and brought them into a structured QGIS project.

The case study covers 656 restaurants, 77,302 buildings, 11,992 road objects, 3 inundation scenarios, 3 evacuation routes and 3 risk states for baseline and climate development to 2050.

Domain results from the reference run

Climate shift and inundation impacts.

These values belong exclusively to the documented reference run and do not change when comparison runs are added.

Shares of the four risk classes for baseline, RCP 4.5 and RCP 8.5 in the documented reference run.
Shift in climate-risk classesShares of the four risk classes for baseline, RCP 4.5 and RCP 8.5 in the documented reference run.
Affected area, buildings and road length at 1, 2 and 5 metres in the documented reference run.
Impact of inundation depthsAffected area, buildings and road length at 1, 2 and 5 metres in the documented reference run.

Inundation screening

Depthkm²BuildingsRoad km
1 m11.2385,741121.68
2 m13.6067,914164.45
5 m20.01213,394279.62

Risk classes to 2050

ScenarioLowMediumHighVery High
Baseline2.60 %75.90 %21.30 %0.20 %
RCP 4.50.30 %13.70 %84.50 %1.50 %
RCP 8.50.00 %15.70 %83.10 %1.20 %

Three evacuation routes

Startmminvulnerable kmvulnerable
Santini1,975.723.70.0010.1 %
Restaurant Spvgg Cannstatt1,933.123.20.0010.1 %
Köfteci Tuncay1,975.123.70.0010.1 %

Inundated area, affected buildings and road length increase monotonically with depth. At the same time, the combined High and Very High share rises from 21.5% at baseline to 86.0% under RCP 4.5; under RCP 8.5 it stands at 84.3%.

Seven charts

Relationships, profiles and climate shift.

The correlation matrix, radar profiles and class shift carry the central findings. Four further charts document distributions and relationships.

Correlation matrix for the six hazard indices with annotated Pearson coefficients.
Correlation matrixCorrelation matrix for the six hazard indices with annotated Pearson coefficients.
Chart as text

The matrix shows which hazards coincide spatially and which run opposite — the basis for weighting the composite.

Six histograms show the distributions of wind, frost, flood, heat, landslide and erosion values.
Hazard distributionsSix histograms show the distributions of wind, frost, flood, heat, landslide and erosion values.
Chart as text

The distributions show which indices vary broadly and which concentrate in particular value ranges.

Radar profiles compare the six hazard components of the five highest-scoring restaurants.
Profiles of the five highest restaurant scoresRadar profiles compare the six hazard components of the five highest-scoring restaurants.
Chart as text

The profiles show that similarly high composite scores can arise from different combinations of flood, erosion, heat and other hazards.

Bar chart showing restaurant counts in four risk classes.
Restaurants by risk classBar chart showing restaurant counts in four risk classes.
Chart as text

At baseline, 2.6% are low, 75.9% medium, 21.3% high and 0.2% very high.

Scatter plot of terrain elevation and flood index with a trend line.
Elevation and flood scoreScatter plot of terrain elevation and flood index with a trend line.
Chart as text

The chart tests the expected relationship between low elevation and higher flood score while showing the variation introduced by other components.

Box plots compare all six hazard indices across the four restaurant risk classes.
Hazards by risk classBox plots compare all six hazard indices across the four restaurant risk classes.
Chart as text

The distributions show which hazards rise with the composite class and how strongly the classes overlap.

Bar chart comparing restaurant risk classes for baseline, RCP 4.5 and RCP 8.5.
Climate-driven class shiftBar chart comparing restaurant risk classes for baseline, RCP 4.5 and RCP 8.5.
Chart as text

The combined High and Very High share rises from 21.5% at baseline to 86.0% under RCP 4.5; under RCP 8.5 it stands at 84.3%.

Maps and QGIS views

Six views lead through the spatial result.

Project setup, composite, building exposure, inundation, evacuation and climate comparison form the primary map sequence.

DEM, local OSM basemap, restaurants, buildings and districts in the Stuttgart study area
Project overview and data basisDEM, local OSM basemap, restaurants, buildings and districts in the Stuttgart study area
The weighted composite combines valley, water, heat and slope patterns with restaurant points
Composite hazard mapThe weighted composite combines valley, water, heat and slope patterns with restaurant points
77,257 building polygons show elevated classes mainly in lower corridors
Buildings by flood risk77,257 building polygons show elevated classes mainly in lower corridors
Nested 1 m, 2 m and 5 m zones follow waterways and lower terrain
Inundation zonesNested 1 m, 2 m and 5 m zones follow waterways and lower terrain
Three pedestrian routes lead from high-scoring restaurants to a safe target, with vulnerable segments highlighted
Evacuation routesThree pedestrian routes lead from high-scoring restaurants to a safe target, with vulnerable segments highlighted
The comparison shows the increase in High and Very High restaurant classes by 2050
Baseline and RCP 8.5The comparison shows the increase in High and Very High restaurant classes by 2050

Domain context

Scope and limitations of the analysis.

The limitations are part of interpreting the maps and metrics and form part of the documented result.

  1. 01

    The 30 m raster smooths small terrain features.

  2. 02

    The reference run uses a DEM-based HAND approximation, not a hydraulic simulation.

  3. 03

    OpenStreetMap completeness and detail vary spatially.

  4. 04

    The climate factors represent a regional 2050 scenario, not a local extreme-event simulation.

  5. 05

    The results are a comparative spatial screening, not an official flood or hazard map.

Generated geodata

From the 30 m raster to evacuation routes.

The inventory separates input and context data, terrain derivatives, hazard products and result vectors. The download package contains a file-level manifest.

Input and context

DatasetGeometryFormatCRSExtentCountSource and role
Study areaextentQGIS / GeoPackageEPSG:2583210 × 10 km1OpenStreetMap Nominatim
shared analysis extent
Elevation modelrasterGeoTIFFEPSG:25832342 × 345 cells117,990Copernicus DEM GLO-30
30 m basis for terrain and hazard analysis
RestaurantspointsGeoPackage / GeoJSON / CSVEPSG:2583210 × 10 km656OpenStreetMap
point exposure analysis
BuildingspolygonsGeoPackageEPSG:2583210 × 10 km77,302OpenStreetMap
polygon flood exposure
Drivable roadslinesGeoPackageEPSG:2583210 × 10 km11,992OpenStreetMap
line flood exposure
Pedestrian networkedgesOpenStreetMap graphEPSG:2583210 × 10 km94,304OpenStreetMap
evacuation-route calculation
City districtspolygonsGeoPackageEPSG:25832Stuttgart19OpenStreetMap
spatial aggregation

Terrain derivatives

DatasetGeometryFormatCRSExtentCountSource and role
Terrain rastersrastersGeoTIFFEPSG:25832342 × 345 cells each9Copernicus DEM derived
slope, aspect, TPI, TWI, TRI, profile curvature, flow accumulation, HAND and hillshade

Hazards and climate

DatasetGeometryFormatCRSExtentCountSource and role
Six hazard indicesrastersGeoTIFFEPSG:25832342 × 345 cells each6terrain and urban derivatives
wind, frost, flood, heat, landslide and erosion on [0,1]
Multi-hazard compositerasterGeoTIFFEPSG:25832342 × 345 cells1weighted hazard indices
combined spatial risk representation
Inundation zonespolygonsGeoPackageEPSG:2583210 × 10 km3 scenariosDEM-based HAND approximation
screening for 1 m, 2 m and 5 m
Climate scenariospoints and tablesGeoPackage / CSVEPSG:25832656 restaurants3 × 656EURO-CORDEX orientation factors in the task
baseline, RCP 4.5 and RCP 8.5 to 2050

Result vectors

DatasetGeometryFormatCRSExtentCountSource and role
Restaurants with risk attributespointsGeoPackage / GeoJSON / CSVEPSG:2583210 × 10 km656sampled hazard rasters
six hazards, composite and risk class
Buildings with flood riskpolygonsGeoPackageEPSG:2583210 × 10 km77,257sampled flood raster
building exposure in four classes
Roads with flood risklinesGeoPackageEPSG:2583210 × 10 km11,992sampled flood raster
road exposure in four classes
High-risk selectionpointsGeoPackageEPSG:2583210 × 10 km33restaurant composite
restaurants with composite above 0.50
Evacuation routeslines and pointsGeoPackageEPSG:25832central Stuttgart3OpenStreetMap pedestrian network
routes, starts and targets with vulnerable segments

QGIS project with geodata

Open the result locally in QGIS.

The curated package contains the portable project, referenced data and a file-level manifest.

ZIP · 27.7 MB

Download QGIS package
  • relative QGIS data sources
  • GeoTIFF, GeoPackage, GeoJSON and CSV
  • German and English README files
  • data manifest and attribution

Contains information from OpenStreetMap: © OpenStreetMap contributors, ODbL 1.0. The elevation model and derivatives use Copernicus DEM GLO-30. Complete mappings are included in the package.

OpenStreetMap licence ↗ · Copernicus DEM ↗ · Public benchmark data ↗