Case Study · Scientific & Environmental Software
NOAA Phytoplankton Monitoring Network
An operational scientific data platform integrating more than two decades of observations from multiple generations of collection systems into a standardized, quality-controlled archival workflow.
PMN collects water-quality and phytoplankton observations from volunteer monitoring groups across the United States. These records support harmful algal bloom monitoring, long-term environmental assessment, and public access to coastal ecosystem data.
- 23+ Years
- historical observations
- 5 Sources
- integrated systems
- 10 QA Flags
- added to the archive
- 2 Modes
- incremental sync + complete rebuild
The Challenge
One archive. Multiple generations of data.
More than 23 years of observations had accumulated across NOAA databases, archival datasets, historical spreadsheets, third-party submissions, and regional Google Forms. Each source had evolved with different schemas, taxonomic conventions, validation rules, and quality-control practices.
The platform needed to reconcile those sources into a single authoritative archive, preserve historical records, identify new observations, and apply consistent scientific quality control. Its output also needed to fit NOAA’s operational publication environment.
System Architecture
From source records to an operational archive.
Five ingestion sources—ERDDAP, NCEI, Google Forms, local files, and Ocean & Earth—feed a common processing workflow. Structured metadata and configuration guide the transformations.
Harmonize source schemas
Use structured configuration for source-specific schemas, field mappings, missing-value conventions, and transformation rules.
Clean and reconcile records
Standardize and deduplicate historical and current observations, preserving historical records while identifying new observations.
Enrich taxonomy
Apply metadata-driven taxonomic mappings to reconcile conventions across generations of collection systems.
Validate and flag
Apply IOOS QARTOD testing and additional spatial validation, generating standardized QA flags before publication.
ERDDAP-compatible output → NOAA ERDDAP archive
Automated archival publication brings the standardized, enriched, and quality-controlled dataset back into ERDDAP as its final archival destination.
Automated QA/QC
Consistent quality flags across the archive.
Ten standardized QA flags were added to the archive, covering the following fields:
- Air temperature
- Latitude
- Longitude
- Salinity
- Sample site
- Water temperature
- Cell count
- pH
- Dissolved oxygen
- Windspeed
Operational Reliability
Designed for routine updates, full rebuilds, and easier troubleshooting.
I designed the platform so NOAA could operate the archive repeatedly over time, not just produce a one-time cleaned dataset. The workflow supports both routine ingestion of new observations and complete reprocessing of the historical archive, while providing enough visibility to understand what happened during each run.
Make each run easier to inspect
Reliable operation also requires visibility into what the pipeline is doing. I added several mechanisms to make processing easier to verify and troubleshoot:
- Logs and diagnostics
- Capture processing activity and provide a starting point for investigating failures.
- Intermediate outputs
- Expose results from key processing stages so transformed data can be inspected before final publication.
- Record accounting by source
- Tracks record counts across individual source systems, making it easier to identify missing, duplicated, or unexpectedly changed data.
Together, these capabilities turned the workflow from a collection of manual data-processing steps into a repeatable operational system that reduced the effort required to aggregate, clean, validate, and republish PMN data.
Technical Ownership
Scientific requirements through implementation.
Primary development and technical ownership included working directly with NOAA scientists, designing the end-to-end ingestion and processing workflow, and implementing source-specific transformation, standardization, and quality-control logic.
The work also covered metadata-driven schema and taxonomic mappings, deployment support, troubleshooting, and long-term maintenance.
Start a Conversation
Bring your scientific data into reliable operation.
Tell us about your data sources, your current workflow, and what needs to work better.