Skip to content

[Feature] Support richer data types, starting with VECTOR<T, N> based on PIP-40 #197

Description

@ChaomingZhangCN

Search before asking

  • I searched in the issues and found nothing similar.

Motivation

Apache Paimon Java has introduced the VECTOR<T, N> data type based on PIP-40. It supports schema representation, regular data-file storage such as Parquet, dedicated vector storage, reads, writes, and Data Evolution.

Paimon C++ currently recognizes .vector. file names for some file-level bookkeeping, but it does not yet provide a VECTOR logical type, Arrow mapping, serialization, storage, or end-to-end read and write support.

This issue implements the VECTOR roadmap item tracked in #186.

Solution

Introduce VECTOR<T, N> support incrementally, while keeping the schema and storage behavior compatible with Apache Paimon Java.

Phase 1: Schema and regular Parquet storage

  • Add VECTOR<T, N> to the Paimon C++ logical type system.
  • Implement schema JSON serialization and deserialization compatible with Paimon Java.
  • Map VECTOR<T, N> to Arrow FixedSizeList<T, N>.
  • Support the element types defined by PIP-40:
    • BOOLEAN
    • TINYINT
    • SMALLINT
    • INT
    • BIGINT
    • FLOAT
    • DOUBLE
  • Validate that the dimension is positive and fixed.
  • Validate that the written vector length equals N.
  • Reject null vector elements.
  • Support reading and writing VECTOR columns in regular Parquet data files.
  • Add end-to-end append-table tests.
  • Add Java/C++ schema and file compatibility tests.

Phase 2: Schema Evolution, Data Evolution, and Primary-Key Table Baseline

Schema Evolution
  • Support adding and dropping VECTOR columns.
  • Support reading files written with previous table schemas.
  • Preserve VECTOR values when projecting files across schema versions.
  • Reject incompatible dimension changes, such as
    VECTOR<FLOAT, 3> to VECTOR<FLOAT, 5>.
  • Reject incompatible element-type changes.
  • Add schema-evolution integration tests for both append-only and
    supported primary-key tables.
Data Evolution Read/Write
  • Support VECTOR columns in row-tracking append-only tables with
    data-evolution.enabled = true.
  • Support full-row writes containing VECTOR columns.
  • Support partial-column writes containing VECTOR columns through the
    write-schema path.
  • Merge VECTOR values from files covering the same row-id range.
  • Support reading VECTOR columns across different file schema IDs.
  • Add Data Evolution integration tests covering full writes,
    partial VECTOR writes, null values, and mixed old/new schema files.
Primary-Key Table Baseline
  • Allow VECTOR columns as ordinary non-key value columns in
    deduplicate primary-key tables backed by regular Parquet files.
  • Support insert, same-key update, and read.
  • Support write-buffer spill and compaction.
  • Support adding and dropping VECTOR value columns while retaining
    readability of files written with previous schemas.
  • Reject VECTOR columns as primary, partition, bucket, sequence,
    sequence-group ordering, or sorting fields.
  • Explicitly reject lookup, partial-update, aggregation, and other
    unsupported primary-key configurations until they are implemented.
  • Add end-to-end primary-key table tests for write, update, read,
    spill, compaction, and schema evolution.

Phase 3: Dedicated vector storage

  • Support the vector file format configuration used by Paimon Java.
  • Support reading and writing dedicated *.vector.vortex files.
  • Integrate vector files with row tracking and Data Evolution.
  • Integrate vector files with scan planning, file commits, and conflict handling.
  • Add Java, Python, and C++ Vortex compatibility tests.

Initial scope

The initial implementation can focus on Phase 1, providing a usable end-to-end vertical slice through schema representation, Arrow mapping, and regular Parquet reads and writes.

Dedicated Vortex vector storage and Data Evolution can be delivered through follow-up pull requests under this issue.

The following items are not required for the initial implementation:

  • ORC VECTOR support
  • Vector indexes or similarity search
  • Changing vector dimensions through schema evolution
  • VECTOR values inside shared-shredding MAP columns
  • Element types not supported by Apache Paimon Java

PIP-40 should only be considered fully supported after all phases are complete. Completing Phase 1 means that regular Parquet VECTOR storage is supported, but does not imply support for dedicated Vortex vector files.

Anything else?

Related roadmap: #186

Are you willing to submit a PR?

  • I'm willing to submit a PR!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions