Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

44 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

improv — image provenance

improv is a shared data platform for scientific imaging instruments. It stores, organizes, and queries images and their associated scientific products — regardless of instrument type or scale.

Every image accumulates an append-only provenance log: geolocation, segmentation outputs, classifier scores, human annotations, oceanographic context. Records are never deleted or overwritten. A classifier can be re-run years later and its outputs attach to the same images alongside the original run.

Architecture

improv has three storage layers:

  • Object store — raw image bytes and binary products (segmentation masks, etc.), keyed by image ID
  • Columnar store — queryable image metadata, provenance records, and plugin index tables (DuckDB+Parquet or VAST DB)
  • OLTP database — mutable organizing metadata: instruments, samples, datasets, ingest tasks (PostgreSQL or SQLite)

A plugin system extends provenance handling. Plugins register via dependency injection — the core never imports them at load time. Each plugin handles a specific provenance kind and optionally maintains an index table for fast querying. Plugins are generic and parameterized (geolocation, sample context, machine classification); instrument-specific presets live under improv.plugins.ifcb (IFCB morphometric features, IFCB CNN classification), pinning a kind/index_table onto a generic plugin.

Access patterns

  • By time and instrument
  • By spatial bounding box (lat/lon/depth)
  • By named dataset (defined as time spans)
  • By sample (for discrete-sample instruments)
  • By provenance kind

REST API

improv exposes a FastAPI service with endpoints for image data and metadata, provenance, instruments, samples, datasets, and ingest task tracking. Classifier support adds taxonomy registration/lookup and two decode paths: a stateless decode (caller supplies a vector) and a decoded read that fetches an image's classification provenance and resolves each record against its own model_version. A thin HTTP client (improv.client.ImprovClient) is provided for ingest scripts that need OLTP access without direct database credentials, including taxonomy registration.

Ingest architecture

Batch producers (ingest pipelines, classifiers) use a hybrid approach:

  • OLTP operations (register instruments, samples, ingest tasks, classifier taxonomies) go through the REST API
  • High-volume writes (image metadata, provenance, index tables, image bytes) go directly to the columnar store and object store

This avoids coupling ingest scripts to the database while keeping high-throughput writes off the HTTP path. Low-volume producers that prefer not to depend on the columnar store directly can post small provenance batches over REST (POST /images/provenance/batch, one instrument per batch).

Idempotency

The provenance log is append-only, so writes are never overwritten — idempotency is achieved by appending, then deduplicating at read time. This works identically on every backend (VAST DB, DuckDB+Parquet, and any future one), because append is the only operation they all share.

  • A provenance record is identified by (image_id, kind, source, data_hash), where data_hash is a canonical RFC 8785 (JCS) hash of the data payload, stamped server-side. Re-posting a byte-identical record re-appends a row that collapses to one on read; a genuinely different payload (e.g. a new model_version) hashes differently and is retained. Canonicalization normalizes key order and number formatting (1.0 == 1) across producers in different languages, and rejects NaN/Infinity.
  • Index records are deterministic projections deduplicated on their full column tuple.

Client contract: a record's timestamp is the event time (when the image was collected, or the classifier result produced) — a property of the observed fact, captured once. It is not the time of the HTTP request. Retries must resend the identical record, so put real event time in timestamp, never wall-clock-at-send; otherwise each attempt looks distinct and will not deduplicate. (The server separately records its own write time.)

Install

pip install .            # base — columnar store, object store, models, client
pip install '.[db]'      # adds SQLAlchemy for direct OLTP access
pip install '.[service]' # adds FastAPI, CLI, migrations

Dependencies

Package Role
amplify-db-utils Columnar storage (DuckDB+Parquet / VAST DB)
amplify-storage-utils Object storage (HashdirStore / S3)
pydantic Models and validation
pyarrow Columnar data exchange
rfc8785 Canonical (JCS) hashing of provenance payloads for idempotency
httpx Thin ingest client
fastapi, sqlalchemy, alembic Service extras

About

a shared data platform for scientific imaging instruments. It provides a single place to store, organize, and query images and their associated scientific products — regardless of instrument type or scale.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages