Case Study · Scientific & Environmental Software

NOAA Phytoplankton Monitoring Network

An operational scientific data platform integrating more than two decades of observations from multiple generations of collection systems into a standardized, quality-controlled archival workflow.

PMN collects water-quality and phytoplankton observations from volunteer monitoring groups across the United States. These records support harmful algal bloom monitoring, long-term environmental assessment, and public access to coastal ecosystem data.

23+ Years
historical observations
5 Sources
integrated systems
10 QA Flags
added to the archive
2 Modes
incremental sync + complete rebuild

The Challenge

One archive. Multiple generations of data.

More than 23 years of observations had accumulated across NOAA databases, archival datasets, historical spreadsheets, third-party submissions, and regional Google Forms. Each source had evolved with different schemas, taxonomic conventions, validation rules, and quality-control practices.

The platform needed to reconcile those sources into a single authoritative archive, preserve historical records, identify new observations, and apply consistent scientific quality control. Its output also needed to fit NOAA’s operational publication environment.

System Architecture

From source records to an operational archive.

Five ingestion sources—ERDDAP, NCEI, Google Forms, local files, and Ocean & Earth—feed a common processing workflow. Structured metadata and configuration guide the transformations.

  1. Harmonize source schemas

    Use structured configuration for source-specific schemas, field mappings, missing-value conventions, and transformation rules.

  2. Clean and reconcile records

    Standardize and deduplicate historical and current observations, preserving historical records while identifying new observations.

  3. Enrich taxonomy

    Apply metadata-driven taxonomic mappings to reconcile conventions across generations of collection systems.

  4. Validate and flag

    Apply IOOS QARTOD testing and additional spatial validation, generating standardized QA flags before publication.

ERDDAP-compatible output → NOAA ERDDAP archive

Automated archival publication brings the standardized, enriched, and quality-controlled dataset back into ERDDAP as its final archival destination.

Automated QA/QC

Consistent quality flags across the archive.

Ten standardized QA flags were added to the archive, covering the following fields:

  • Air temperature
  • Latitude
  • Longitude
  • Salinity
  • Sample site
  • Water temperature
  • Cell count
  • pH
  • Dissolved oxygen
  • Windspeed

Operational Reliability

Designed for routine updates, full rebuilds, and easier troubleshooting.

I designed the platform so NOAA could operate the archive repeatedly over time, not just produce a one-time cleaned dataset. The workflow supports both routine ingestion of new observations and complete reprocessing of the historical archive, while providing enough visibility to understand what happened during each run.

Keep the archive current

Incremental synchronization

Used for recurring updates as new monitoring observations become available.

The pipeline identifies new records, processes them through the same standardization and quality-control workflow used by the archive, and generates ERDDAP-compatible output for publication. This allows the dataset to grow without repeatedly rebuilding the full historical record.

Rebuild the archive when requirements change

Complete historical rebuild

Used when a change needs to be applied consistently across the entire dataset.

The platform can reprocess more than 23 years of observations through the full integration, enrichment, and quality-control workflow and regenerate the ERDDAP-compatible archival dataset. This makes it possible to apply updated schemas, metadata, taxonomic information, or QA/QC logic across both new and historical records.

Make each run easier to inspect

Reliable operation also requires visibility into what the pipeline is doing. I added several mechanisms to make processing easier to verify and troubleshoot:

Logs and diagnostics
Capture processing activity and provide a starting point for investigating failures.
Intermediate outputs
Expose results from key processing stages so transformed data can be inspected before final publication.
Record accounting by source
Tracks record counts across individual source systems, making it easier to identify missing, duplicated, or unexpectedly changed data.

Together, these capabilities turned the workflow from a collection of manual data-processing steps into a repeatable operational system that reduced the effort required to aggregate, clean, validate, and republish PMN data.

Technical Ownership

Scientific requirements through implementation.

Primary development and technical ownership included working directly with NOAA scientists, designing the end-to-end ingestion and processing workflow, and implementing source-specific transformation, standardization, and quality-control logic.

The work also covered metadata-driven schema and taxonomic mappings, deployment support, troubleshooting, and long-term maintenance.

Start a Conversation

Bring your scientific data into reliable operation.

Tell us about your data sources, your current workflow, and what needs to work better.