Some freewriting dumb dumb
Introduce real world problem Nitrate in the water is bad. Iowa has a lot of nitrate in the water. It mostly comes from fertilizer on corn fields. It varies seasonally. Long-term exposure
Clearly state team’s research question State and federal agencies deploy real-time nitrate sensors in Iowa’s waterways. Coverage is spotty – sensors malfunction frequently and aren’t deployed everywhere. More recently, many of the water-monitoring programs have been cut, and so data collection is getting even harder. We wondered: could we create virtual sensors which could predict nitrate levels in any part of the state from other more robust data sources?
Steps to improve project Point pollution sources are potentially a huge missing piece of our pipeline. Latitude and longitude are rough proxies for it, but it would be useful to include them as it would likely improve performance. We didn’t include it in this project due to time constraints, statewide point pollution data sources turned out to be prohibitively difficult to gather in the allotted time.
Improved land use data
Our models would also be improved by better network modeling.
Draft
The Problem Setup
Iowa grows a lot of corn, 2.5 Billion bushels in 2024, or enough to feed the entire United States for a year. Iowa is also the only state in the US whose yearly cancer rate is increasing, a trend that began back in 2015. These two things are linked.
You see, the soil can’t natively handle that kind of corn production – it requires a lot of extra nitrogen, added in the form of artificial fertilizer and animal manuer – and not all of that nitrogen stays in the soil. The excess leaches into streams, rivers and eventually the manucipal watersupply as nitrate, a carcinogen. The typical background concentatration of nitrate in water is 1 to 2 mg/L, but Iowa is unique in that its residential water frequently exceeds a nitrate concentration of over 10 mg/L, the federally mandated limit.
This is why both state and federal agencies deploy real-time sensors in Iowa’s waterways to monitor the nitrate levels. However, coverage is spotty – sensors break, are expensive to deploy, and many have recently been retired due to (potentially politically motivated) budget cuts. This leaves residents, policy makers and farmers flying blind.
Research Question
Luckily, the pathway nitrate takes from the farm to your brita filter is well understood: water flows through the farms into streams and rivers, carrying excess nitrate with it. We wondered, can we predict nitrate risk at a specific location using ONLY publicly available weather and land-use data, without a physical sensor? In short: yes, we can.
The data
We started with a list of 162 real sensors deployed across Iowa’s waterways, and after filtering out for large data gaps and malformed data, we kept the 85 best, with records spanning 4 to 15 year periods betwen 2008 to 2025. Here’s an example of one of them, and this is what it’s nitrate data looks like.
From these, we built two outcomes: (1) the daily maximum nitrate concentration, a regression target, and (2) a daily yes/no violation flag answering the question: on this day, did the sensor observe a nitrate level above the federal safety limit?
Next, for each sensor, we computed its drainage basin, the area of land which actually flows into the sensor. It was vital we get this step correct: incorrectly including land would pollute models with irrelevant data, and incorrectly excluding land would deprive models of valuable signal. So we computed the basin for each site in three different ways, and then manually reviewed each to ensure its validity.
We then layered four datasets onto each of these drainage basins: annual crop distribution data from satallite imagery, annual surplus nitrogen soil data obtained from a nature paper and two different source of daily historical weather data containing things like rainfall, temperature, and evapotranspiration information. These came at wildly different spatial resolutions, so we stitched these together to a common grid, which you see here. We also computed basic geographic information for each of these grid cells, including the area of the cell contained in the basin and the downhill flow distance from the center of each cell to the sensor. The predictors which carried the most weight in our models turned out to be the crop mix near the sensor, recent rainfall, and most importantly the location of the sensor itself.
Modeling approach
We used gradient-boosted decision trees, XGBoost, for both of our models. The relationship between weather, fertilizer and water-borne nitrate is highly nonlinear and the data (especially the sensor data) is messy, so trees decisevly beat classical time-series forecasters in every comparison we tried.
It was also important to be careful about our validation strategy. If sensor A is a few kilometers downstream of sensor B, then the two sites will report highly correlated nitrate values. Since our goal was to predict nitrate at unseen locations, we had to ensure that sensors with overlapping basins were held out on the same side of the train/test split to prevent data leakage and ensure our reported accuracy reflected the true generalization to new places. Basically, we needed to make sure our models actually learned to predict nitrate from weather and crop data, and weren’t simply memorizing patterns from location to location.
Findings
A few things stood out. First, nitrate has a strong memory, so the best predictor of today’s nitrate value is yesterday’s nitrate value, although this of course this baseline only makes sense when you have a sensor at your prediction location. Second, nitrate is both incredibly seasonal – it consistently peaks in the late spring and early summer all across the entire state – and is highly variable by location, with some extreme outlier states reporting an average nitrate concentration at a whopping 35mg/L.
The exciting result is classifiation: our current best model can flag a violation risk at around INSERT ACCURACY. With a bit of tuning, we can catch… but this is already a useful signal in places lacking a sensor.
Predicting exact concentration is much harder, <quantify results>. However this is still a positive signal, and the model’s methodology appears to be sensible: land immediately around the sensor contributes far more than the outer reaches of the basin, rainfall from the previous few weeks matters more than rainfall today, consistent with slow movement of nitrate through soil, and there’s a clear seasonal spike in spring/summer.
Next steps
There are two main routes to improving our results: incorporating more features and using more sophisticated modeling techniques.
We expect the best feature we could add is point source pollution data. Iowa has a lot of pig and cow farms, and the manure runoff from those farms are consistent sources of nitrate. We expect latitude and longitude are important feature precisely because they serve as a rough proxy for location dependent factors like this one. Other predictors that could prove useful include soil composition data to explain geographic differences in soil composition and
Theorem
Erin Draft
The Problem Setup
(Isaac) Iowa grows a lot of corn, enough calories for everyone in the United States to eat nothing else, and the soil can’t naturally handle that kind of maize production. It requires a lot of extra nitrogen, added in the form of artificial fertilizer and animal manure, and not all that nitrogen stays in the soil. When it rains, excess nitrogen leachs into streams, rivers and eventually the municipal water supply as nitrate, a carcinogen. The typical background concentration of nitrate is 1 to 2 mg/L of water, but residential water in Iowa frequently exceeds 10 mg/L, the federally mandated limit. This is a big reason why cancer is so prevelant in Iowa, and why since 2015, it is the only US state with a RISING number of annual cancer cases.
(Erin) It’s also why state and federal agencies deploy real-time sensors in Iowa’s waterways to monitor the nitrate levels. However, coverage is spotty – sensors break, are expensive to deploy, and many have been decommissioned following recent budget cuts. This leaves residents, regulators and farmers flying blind in the exact locations they most need information.
But what if it was possible to predict nitrate from other sources of data, to put virtual sensors in the river if you will? Given that nitrate is carried away from fields by runoff, if you know something about where the fertilizer is and what the weather is like, you should be able to say something about the nitrate – that’s what my team thought at least. We wanted to know: using only weather and landuse data, can you flag days with high violation risk?
Luckily, the pathway nitrate takes from the field to to the tap is well understood: water flows through the farms into streams and rivers, carrying excess nitrate with it.We wondered, can we flag nitrate violation risk WITHOUT a physical sensor? In short: yes, we can.
Data Story
What you’re seeing now is the widget my team built to help explore and visualize all the data. We started with 162 sensors, and then filtered down to the 85 best. For each site we then computed the drainage basin in three different ways. For most sites these three basin choices agreed, but for others they were quite different. It was important to get this right, so we manually went through our sites and picked the basins that made the most sense.
We then layered crop, soil-nitrogen surplus and daily weather onto our basins, all stitched toegether on a common grid, which you see here with cells colored by dominant crop type. Here is a graph of the sensor’s nitrate and precipiation data for completeness.
We then spent some time exploring the data we’d curated and found three big things that informed our modeling.
First: water-borne nitrate has a strong memory; if a sensor read a high value yesterday, it will likely read a high value today.
Second: Nitrate is highly seasonal, it climbs in April, peaks in June, and falls in the late summer, coinciding with the growing season of corn.
Third: Sensing sites have high variability. Most are pretty small, but some are truly massive. Some sites almost never see a nitrogen violation, others stay consistently above the 10 mg/L threshold.
Second: Key EDA insight: nitrate has strong memory (a persistence baseline beat linear regression), clear seasonality beyond what rain explains, and huge variation by location — some sites average over 30 mg/L. Main limitation: sensor coverage is sparse, and annual crop/soil data misses within-season change.
Final data list
Our finalized feature list consisted of
- daily weather data, stuff like precipitation, temperature and solar radiation
- annual crop data
- annual nitrogen soil surplus data
- geographical features like latitude and longitude
- calendar information to help with seasonality adjustments
- and average statewide nitrate values, giving us
- 68 total features.
We found that grid cells near the sensor mattered far more than grid cells far away, so accounted for this by aggregating data into near and far distance buckets and by directly weighting values by distance.
Modeling Journey (~40s)
We tried linear/logistic regression and traditional timeseries modeling experiments but all of these failed miserably. The first positive result we got was from XGBoost, a gradient boosted tree-ensemble model, whose baseline results at least gave us positive predictions for both regression and classification targets.
Being honest about our validation was critical: as you can see, basins nest insides of each other, and sensors with nested basins report correlated values. To ensure we didn’t leak our target values to the models during training, we kept overlapping families of basins on the same side of the test/training split.
Final Results (~55s)
Our regression results are modest, but we still get positive predictive power; our model explains about a third of daily variance on unseen basins.
Flagging violations works much better. This score is the one we value the most, it tells us roughly how well our model balances true positives with false alarms, and our results are 2.4 x better than a “random guess” baseline. With proper tuning, we can do even better.
To see what I mean, let’s look at the deployed model in our widget. We drop a pin anywhere in Iowa, and it produces a basin for us. We can then run a forecast on a chosen model year, and we get out predicted regression and violation probability curves.
By changing this parameter, we can change the tradeoff between the number of true positives we detect and the number of false alarms we’re willing to tolerate. With it set to 2, we expect to catch 86% of true positives but will have to tolerate over 50% of our violation flags being false alarms.
Predicting exact concentration is harder, as expected: on unseen basins the regressor explains about a third of the daily variance (R² ≈ 0.33), with typical error near 4 mg/L. The decomposition is the real story — it’s twice as good at tracking movement within a basin (within-site R² ≈ 0.42) as at ranking the absolute level of a brand-new basin (between-site R² ≈ 0.21). So it’s a strong “when and roughly how bad,” not a calibrated meter.
[Visual: spatial-CV precision–recall curve. Visual: predicted-vs-actual error band.]
The model’s behavior is sensible: land near the sensor matters more than distant basin area, rain lag matters more than same-day rain, and there’s a clear spring/early-summer spike. Both gain and permutation importance agree on the backbone — the statewide nitrate level that day, the sensor’s location and distance geometry, and near-field corn — and permutation confirms the model leans on what’s immediately upstream, exactly the riparian effect we engineered for.
Limitations & Next Steps (~35s)
Business / Decision Relevance (~30s)
Final Transcript
Introduction (currently 60s)
(Isaac) Iowa grows a lot of corn – enough calories per year for everyone in the United States to eat nothing else. This is only possible by adding a huge amount of extra nitrogen to the soil, however, not all that nitrogen ends up in crops. It leaches into streams, rivers, and eventually the municipal water supply, as nitrate, a carcinogen. The typical background rate of nitrate is around 1 mg/L, but tap water in Iowa frequently exceeds the federal limit of 10 mg/L, which itself is 3x higher than the threshold for increased cancer risk. This is why Iowa is the only US state with a growing cancer rate.
(Erin) It’s also why state and federal agencies use real time sensors to monitor nitrate in Iowa’s waterways. However, coverage is spotty – sensors break and are not longer being consistently replaced due to budget cuts. This leaves utilities, farmers and regulators flying blind in large swathes of the state.
(Preet) But what if you could predict nitrate from other, more consistent data – put a “virtual sensor” in the water if you will? Our team thought this might be possible. We asked: if you know what’s planted where, and what the weather has recently been like, can you flag days at a specific location with high nitrate violation risk?
Data Setup (currently 40s)
(Isaac) What you’re seeing now is the widget we created to view and explore the data. We started with 162 sensors and cut this down to the 85 best.
(Erin) We then computed the drainage basin of each site in three different ways. For many sites these three choices matched quite well, but for others they were wildly different. Getting this correct was critical for everything downstream in the project, so we manually selected the correct basin for every site.
(Preet) We then layered annual land-use and daily weather data onto a common grid, which you see here colored by dominant crop type. Here’s a graph showing the daily rain and nitrate values for this site, the purple line is the nitrate signal we’re trying to model.
EDA (currently 30s)
(Isaac) We then spent some time exploring the data we currated and found three big things that informed our modeling.
First, waterborne nitrate has a strong memory; if a site had high nitrate yesterday, it probably has high nitrate today.
(Erin) Second, nitrate it highly seasonal, even after controlling for weather: it rises in spring, peaks in June, and falls in late summer – coinciding with the growing season of corn.
(Preet) Third, sensing sites have high variability. Most are placed on small local streams, but some basins are truly massive. Some sites barely ever see a nitrate violation others stay consistently above the 10 mg/L threshold.
Final feature list (currently 30s)
(Isaac) The full list of data we used to build our virtual sensors model consisted of
- daily weather data,
- annual crop and soil nitrogen data,
- geographical features like latitude and longitude
- calendar information to account for seasonality
- average statewide nitrate values giving us
- 68 total features.
(Isaac or maybe Erin take over) We found that grid cells near the sensors matttered far more than grid cells far away. We accounted for this by aggregating data into near and far distance buckets and by weighting values directly by distance (both of which improved our results).
Modeling (currently 35 seconds)
(Erin) We tried both standard regression and timeseries modeling experiments, but these failed miserably and we quickly moved on. They were also poorly suited to our stated goal: we wished to train a model on our 85 physical sensors which could generalize to completely unseen locations.
(Preet) The first positive result we got came from XGBoost, whose initial results gave us hope that we had real predictive power. (this makes sense because… XGBoost is well-suited to our problem because…) Being careful about our validation was also critical – here you see basins nest inside of each other, and we had to ensure our model wasn’t simply memorizing the collective behavior of basin families. We accounted for this with smart validation techniques, so we can be sure that our final results are true representations of our model’s ability to generalize.
Results (currently 25 seconds)
(Preet, cont.) Our regression results are modest, but we still have predictive power; our model explains roughly 1/3 of daily nitrate variation.
(Isaac) Classifying high-risk days works much better. This is the score we value the most; it tells us how well our model balances true violation days with false alarms, and our results are significantly better than a “random guess” baseline. With proper tuning, we can do even better.
Deployment + tuning
(Isaac, cont.)
To see what I mean, let’s look at our deployed model in the widget. We drop a pin anywhere in Iowa and it produces a basin. We can then run a forecast on a chosen model year and we get our predicted regression and violation probability curves. By changing the beta parameter you see here, we can control the balance between the number of true violations we detect and the number of false alarms we’re willing to tolerate. With beta set to 2, we expect to catch 86% of true positives, but we’ll have to tolerate nearly 60% of our violation flags being false alarms.
Limitations and Improvements (currently 30 sec)
(Erin) This model is obviously limited. Our 85 sensors aren’t enough to fully represent all of Iowa, and you can’t learn everything there is to know about nitrate from daily weather data and annual snapshots of land use alone. Indeed, there are other sources of polution besides agriculture. For instance, manure runoff from livestock operations are also big source of nitrate, and we’d like to add it in the future.
(Preet) There’s also a feature currently present in our data which we barely utilize: the network of rivers which connect basins together. We think some graph-based model architecture could take advantage of this richer structure to improve predictions.
Stakeholders and relevance and end (currently 25 seconds)
(Everyone on screen here) Nitrate predictions let three groups act sooner: utilities can start treatment before a spike hits, farmers can better plan fertilizer application and agencies can target scarce funding toward high-risk areas. Virtual sensor models like this one certainly won’t replace physical detectors – but as an alternative to no information at all, even an imperfect prediction is a major upgrade.