scripts/table-catalog/README.md
This directory contains repeatable client-facing checks for the RustFS S3 Tables Iceberg REST Catalog surface. The goal is to keep S3 Tables compatibility claims grounded in runnable scripts or explicit unsupported entries.
For the release-facing support and limitation matrix, see
docs/architecture/s3-tables-support-matrix.md.
Install the client dependencies:
python3 -m pip install 'pyiceberg[pyarrow]' boto3
Start RustFS locally, then run:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--endpoint http://127.0.0.1:9000 \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--bucket rustfs-s3table-smoke \
--replace \
--cleanup
To persist the successful PyIceberg run as live conformance evidence, provide an output path and the operator-recorded deployment metadata:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--endpoint http://127.0.0.1:9000 \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--bucket rustfs-s3table-smoke \
--replace \
--cleanup \
--rustfs-build rustfs-v1.0.0-beta.8 \
--git-sha "$(git rev-parse HEAD)" \
--catalog-backing durable-strong \
--live-evidence-output /tmp/rustfs-pyiceberg-live-evidence.json
The evidence file uses the schema printed by:
python3 scripts/table-catalog/engine_compatibility.py --print-live-evidence-schema
The smoke helper validates the record before writing it. Missing deployment metadata, mismatched expected and observed status, invalid row counts, or claim promotion beyond the client boundary are treated as evidence failures instead of support-matrix proof.
The smoke test covers:
The default profile uses the canonical RustFS catalog URI:
http://127.0.0.1:9000/iceberg
To exercise the MinIO AIStor-style alias exposed by RustFS:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile rustfs-compat \
--endpoint http://127.0.0.1:9000 \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--bucket rustfs-s3table-smoke \
--replace \
--cleanup
rustfs-compat uses:
catalog URI: http://127.0.0.1:9000/_iceberg
REST signing name: s3tables
If the local deployment still requires the standard S3 signing name for the alias path, override it explicitly:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile rustfs-compat \
--rest-signing-name s3
To verify catalog-vended table credentials, enable server-side credential vending and use the vended credential profile:
RUSTFS_TABLE_CATALOG_CREDENTIAL_VENDING=enabled
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile rustfs-vended-credentials \
--endpoint http://127.0.0.1:9000 \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--bucket rustfs-s3table-smoke \
--replace \
--cleanup
This profile uses the configured principal to create the bucket, enable the table bucket, and create the table. After the table exists, it calls the REST credentials endpoint and reloads the PyIceberg catalog with the returned table-scoped S3 access key, secret key, and session token before append, reload, and scan operations.
Before the PyIceberg append, the profile also checks that the returned credential prefix exactly matches the created table warehouse location after canonical S3 URI normalization, including percent-decoding equivalent path encodings. It then runs a direct S3 data-plane scope probe with the returned temporary credentials:
PutObject, HeadObject, GetObject, and DeleteObject must work inside
the returned table warehouse prefix.PutObject and GetObject to the same bucket outside that prefix must be
rejected.The direct REST catalog probes run by default after the PyIceberg append and scan. For deployments that intentionally expose only the core Iceberg REST Catalog table path, skip those probes explicitly:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--skip-catalog-api-probes
The script can print the current conformance inventories without importing PyIceberg, PyArrow, or boto3:
python3 scripts/table-catalog/pyiceberg_smoke.py --print-client-matrix
python3 scripts/table-catalog/pyiceberg_smoke.py --print-engine-compatibility
python3 scripts/table-catalog/pyiceberg_smoke.py --print-production-failure-coverage
python3 scripts/table-catalog/pyiceberg_smoke.py --print-vendor-profiles
python3 scripts/table-catalog/pyiceberg_smoke.py --print-unsupported-inventory
python3 scripts/table-catalog/pyiceberg_smoke.py --print-production-readiness
python3 scripts/table-catalog/engine_compatibility.py --print-live-evidence-schema
Use these outputs when updating release notes, PR descriptions, or follow-up work items. They are intentionally conservative: PyIceberg and DuckDB have separate automated smoke entrypoints. Spark has a repeatable manual/live harness with pinned client package inputs, generated configuration, generated SQL, expected results, and a CI opt-in gate. Trino, Databend, and Snowflake have generated manual probe inputs, but they remain opt-in and do not promote write or full vendor interoperability claims.
Vendor profiles are machine-readable connection references. They include the
catalog URI shape, warehouse shape, signing name, credential model, namespace
model, pagination model, and explicit not-claimed behavior. They are useful for
building migration docs without turning provider references into compatibility
claims. The vendor_profiles object lists every supported reference template;
the selected_vendor_profile object renders the active profile with the
provided endpoint, account, bucket, and warehouse arguments.
Spark vendor config generation only emits S3-compatible data-plane endpoint,
path-style, and static S3 credential properties for profiles that require them.
RustFS and MinIO AIStor use the supplied endpoint as object storage. Alibaba
OSS Tables uses the documented https://oss-{region}.aliyuncs.com public
S3FileIO endpoint by default. AWS S3 Tables and Cloudflare R2 Data Catalog
reference profiles leave provider object I/O endpoint and credential selection
to the provider-specific Spark/Iceberg runtime configuration instead of
inheriting RustFS local defaults.
The standalone engine helper also prints a vendor compatibility audit. This records the public reference source, catalog path, warehouse shape, signing name, auth model, required validation categories, and explicit not-claimed boundaries for AWS S3 Tables, MinIO AIStor Tables, Cloudflare R2 Data Catalog, and Alibaba OSS Tables. Use it to plan compatibility work; do not treat it as live interoperability evidence.
Generate an AWS S3 Tables-style reference profile:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile aws-s3tables \
--region us-east-1 \
--account-id 123456789012 \
--table-bucket analytics \
--print-vendor-profiles
Generate a Cloudflare R2 Data Catalog-style reference profile:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile cloudflare-r2-data-catalog \
--catalog-uri https://example.account.r2.cloudflarestorage.com/catalog \
--warehouse-name analytics \
--print-vendor-profiles
Generate an Alibaba OSS Tables-style reference profile:
python3 scripts/table-catalog/pyiceberg_smoke.py \
--profile oss-tables \
--endpoint https://cn-hangzhou.oss-tables.aliyuncs.com \
--region cn-hangzhou \
--account-id 123456789012 \
--table-bucket analytics \
--print-vendor-profiles
The standalone engine helper prints the same compatibility matrix and can also generate Spark REST catalog input without importing PyIceberg:
python3 scripts/table-catalog/engine_compatibility.py --print-engine-matrix
python3 scripts/table-catalog/engine_compatibility.py --print-vendor-audit
python3 scripts/table-catalog/engine_compatibility.py --print-spark-config
python3 scripts/table-catalog/engine_compatibility.py \
--profile aws-s3tables \
--region us-east-1 \
--account-id 123456789012 \
--table-bucket analytics \
--print-spark-config
python3 scripts/table-catalog/engine_compatibility.py --print-spark-sql --cleanup
python3 scripts/table-catalog/engine_compatibility.py --print-duckdb-rest-sql
python3 scripts/table-catalog/engine_compatibility.py --print-live-conformance --cleanup
python3 scripts/table-catalog/engine_compatibility.py --print-operations-guide
The production failure helper records the negative coverage required before calling a release production-ready and can generate REST probe steps for a live RustFS endpoint:
python3 scripts/table-catalog/failure_coverage.py --print-failure-matrix
python3 scripts/table-catalog/failure_coverage.py \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--rest-path /iceberg \
--print-failure-probes
python3 scripts/table-catalog/failure_coverage.py \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
--print-disaster-recovery-rehearsal
python3 scripts/table-catalog/failure_coverage.py \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
--writer-count 8 \
--maintenance-worker-count 2 \
--iteration-count 50 \
--print-scale-fault-rehearsal
The generated probe plan covers stale-token commit conflicts, missing metadata object rejection, diagnostics/recovery for finalization gaps, maintenance stale plan rejection, and external catalog sync conflicts. These steps are meant to be run against a prepared live table and should be recorded with the exact RustFS build and client versions used.
--rest-path defaults to /iceberg and generated probe paths include the
mounted catalog prefix, for example /iceberg/v1/{warehouse}/.... Use
--rest-path /_iceberg to generate paths for the compatibility alias.
The smoke test also probes catalog-backed advanced Iceberg surfaces:
main cannot be deleted| Client | Current status | Claim |
|---|---|---|
| PyIceberg | Automated smoke target | create namespace, create table, append, reload, scan, metadata-location, refs, views, maintenance, diagnostics, optional catalog-vended table credentials with exact-prefix data-plane scope probe |
| Spark Iceberg REST catalog | Manual/live harness | pinned Spark and Iceberg package inputs, configuration, SQL, run command, expected row count, and cleanup can be generated for a running RustFS endpoint; CI execution is opt-in |
| Trino Iceberg REST catalog | Manual/live read probe | generated catalog properties and a read-only SELECT probe for a table created by PyIceberg or Spark; no write compatibility claim yet |
| DuckDB Iceberg | Automated smoke target | metadata-location read plus generic REST Catalog single-table DDL, DML, schema evolution, snapshots, /iceberg and /_iceberg signing, fail-closed unsupported boundaries, endpoint-disabled non-atomic multi-table mode, concurrent writers, and PyIceberg cross-read |
| StarRocks Iceberg REST catalog | Documented, not automated | external catalog read-path reference only |
| Databend | Manual/live S3 stage probe | generated S3 stage read probe for table data files; Iceberg REST catalog integration is not claimed |
| Snowflake/Open Catalog integrations | Manual reference probe | generated external volume/catalog SQL template; live RustFS interoperability is not claimed |
Production failure coverage is tracked separately from positive client conformance. Positive smoke tests prove a client can create and use a table; failure probes prove RustFS does not silently advance table state when something goes wrong.
The current failure matrix covers:
Do not promote a failure case from probe-required or load-test-required to
an automated claim until the live probe or stress harness is repeatable and its
RustFS build, client version, and expected response shape are recorded.
The disaster recovery rehearsal plan is a machine-readable operator runbook. It does not mutate state by itself and does not claim automatic repair. It records the REST and S3 probes an operator or CI opt-in job should run against a prepared table when validating recovery behavior.
Generate the rehearsal plan:
python3 scripts/table-catalog/failure_coverage.py \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
--print-disaster-recovery-rehearsal
The plan is gated for CI by:
RUSTFS_TABLE_CATALOG_DR_REHEARSAL=1
The generated phases cover:
loadTableloadTable and table warehouse data-plane policy probesRecord the RustFS build, catalog backing mode, table identifier, metadata location, and expected response status for each run. Treat migration blockers, manual-review diagnostics, stale rollback/import conflicts, and data-plane policy failures as fail-closed results that require operator investigation before cutover or release claims.
The scale and fault rehearsal plan is a machine-readable opt-in runbook for production-style stress and failure evidence. It does not run the stress test by itself and does not promote compatibility or scale claims without recorded live results.
Generate the rehearsal plan:
python3 scripts/table-catalog/failure_coverage.py \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--table-warehouse-location s3://rustfs-s3table-smoke/tables/table-id \
--writer-count 8 \
--maintenance-worker-count 2 \
--iteration-count 50 \
--catalog-backing durable-strong \
--print-scale-fault-rehearsal
The plan is gated for CI by:
RUSTFS_TABLE_CATALOG_SCALE_FAULT_REHEARSAL=1
The generated phases cover:
loadTable, table data-plane policy, and operator artifact captureRecord the RustFS build, catalog backing mode, writer count, worker count, iteration count, final metadata location, conflict counts, recovered leases, and failed-closed operations before using the run as release evidence.
The engine helper can print a machine-readable operations guide that ties client conformance, durable backing cutover, maintenance, recovery, permissions, and unsupported-claim governance to the exact evidence operators must record:
python3 scripts/table-catalog/engine_compatibility.py \
--endpoint http://127.0.0.1:9000 \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--print-operations-guide
Use this output as the release checklist when expanding compatibility language. Each section records commands, required evidence, pass criteria, and fail-closed signals. A client or vendor claim should only be promoted when the corresponding live evidence records the RustFS build, catalog backing mode, client version, expected status, observed status, and metadata location.
| Profile | Catalog shape | Warehouse shape | Signing name | RustFS claim |
|---|---|---|---|---|
rustfs | {endpoint}/iceberg | {bucket} | s3 | automated smoke target |
rustfs-compat | {endpoint}/_iceberg | {bucket} | s3tables by default | compatibility smoke target |
rustfs-vended-credentials | {endpoint}/iceberg | {bucket} | s3 | automated credential smoke target when server vending is enabled |
aws-s3tables | https://s3tables.{region}.amazonaws.com/iceberg | arn:aws:s3tables:{region}:{account_id}:bucket/{table_bucket} | s3tables | profile generator only; full AWS S3 Tables API parity is not claimed |
minio-aistor | {endpoint}/_iceberg | {warehouse} | s3tables | profile generator plus RustFS alias smoke; full AIStor extension parity is not claimed |
cloudflare-r2-data-catalog | catalog URI returned by R2 | {warehouse_name} | s3 | profile generator only; live RustFS interoperability is not claimed |
oss-tables | provider REST endpoint | acs:osstables:{region}:{account_id}:bucket/{table_bucket} | osstables | profile generator only; live RustFS interoperability is not claimed |
Unsupported behavior is documented instead of hidden behind internal errors. The current unsupported inventory is:
RustFS advertises table credential scope metadata without returning reusable storage secrets by default. The standard credentials endpoint is registered:
GET /v1/{prefix}/namespaces/{namespace}/tables/{table}/credentials
The endpoint returns an empty storage-credentials list unless table catalog
credential vending is explicitly enabled. LoadTable uses the same issuer when
the request includes X-Iceberg-Access-Delegation: vended-credentials. The
response advertises the issued session for the table warehouse prefix and for
the exact current metadata object; the session policy keeps table data access
inside the warehouse and grants only GetObject to that metadata object.
LoadTable remains metadata-only when delegation is absent, vending is disabled, or the caller lacks the separate table-credentials permission. Disabled and not-authorized fallbacks include an explicit reason. Issuer errors, including disallowed chained temporary credentials, are returned as request errors rather than silently falling back.
The rustfs-vended-credentials profile verifies the client handoff through the
dedicated credentials endpoint. It still uses the configured principal for
setup and REST request signing; the vended credentials are checked against the
created table warehouse location, probed directly against S3 scope boundaries,
and then applied to PyIceberg data-plane access. Stable PyIceberg releases up to
0.11 do not consume LoadTable storage-credentials; native LoadTable coverage
requires a client release with that support.
Enablement is server-side and fail-closed:
RUSTFS_TABLE_CATALOG_CREDENTIAL_VENDING=enabled
RUSTFS_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS=900
The TTL is clamped to the supported short-lived range by the server.
DuckDB can read an individual Iceberg table with iceberg_scan or attach RustFS
as a generic Iceberg REST Catalog. The metadata-location path remains read-only.
The attached catalog path is the prerequisite for DuckDB writes.
Generate the canonical RustFS REST Catalog profile:
python3 scripts/table-catalog/engine_compatibility.py \
--endpoint http://127.0.0.1:9000 \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--rest-path /iceberg \
--rest-signing-name s3 \
--print-duckdb-rest-sql
Generate the compatibility alias profile by changing the last three arguments:
python3 scripts/table-catalog/engine_compatibility.py \
--rest-path /_iceberg \
--rest-signing-name s3tables \
--print-duckdb-rest-sql
The generated ATTACH disables staged create, post-create metadata updates,
multi-table commit, client-side file removal, and purge-on-drop. These options
keep DuckDB within RustFS's claimed single-table REST surface. Do not replace
the explicit endpoint with DuckDB ENDPOINT_TYPE S3_TABLES; that shortcut is
for AWS S3 Tables endpoint and warehouse shapes.
Run the repeatable DuckDB 1.5.5 smoke against an already running RustFS:
python3 scripts/table-catalog/duckdb_smoke.py \
--duckdb /path/to/duckdb \
--endpoint http://127.0.0.1:9000 \
--bucket rustfs-duckdb-smoke \
--namespace duckdb_smoke \
--table events \
--cleanup \
--rustfs-build rustfs-v1.0.0-rc.4 \
--git-sha "$(git rev-parse HEAD)" \
--catalog-backing object \
--live-evidence-output /tmp/rustfs-duckdb-live-evidence.json
The script requires the same PyIceberg, PyArrow, and boto3 dependencies as the
PyIceberg smoke because it verifies both cross-engine directions. It creates an
isolated namespace, keeps the final verified table at two rows for the shared
evidence contract, and cleans all smoke tables only when --cleanup is set.
It refuses to remove pre-existing suffixed smoke tables unless --replace is
set explicitly. Cleanup preserves a namespace that existed before the run.
The automated claim is limited to DuckDB 1.5.5, static S3 credentials, and the
single-table scenarios exercised by this script. It does not claim DuckDB's AWS
S3_TABLES shortcut, staged create, purge-on-drop, format v3, multi-table
atomicity, or catalog-vended credential integration. The smoke verifies that
DuckDB can run a two-table transaction with its multi-table commit endpoint
disabled, but each table remains an independent RustFS commit.
Spark validation should use the same RustFS endpoint and warehouse bucket as the
PyIceberg smoke test. The harness records default pinned client package inputs,
the exact command shape, and the expected row count. It is manual by default and
should only run in CI when explicitly gated with
RUSTFS_TABLE_CATALOG_LIVE_CONFORMANCE=1.
Generate the full live harness document:
python3 scripts/table-catalog/engine_compatibility.py \
--endpoint http://127.0.0.1:9000 \
--warehouse rustfs-s3table-smoke \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--pyiceberg-version 0.10.0 \
--spark-version 3.5.4 \
--iceberg-version 1.7.1 \
--print-live-conformance \
--cleanup
The output includes:
iceberg-spark-runtime and iceberg-aws-bundlespark-sql command using the generated propertiesrow_count=2 before optional cleanupGenerate the configuration properties:
python3 scripts/table-catalog/engine_compatibility.py \
--endpoint http://127.0.0.1:9000 \
--warehouse rustfs-s3table-smoke \
--access-key rustfsadmin \
--secret-key rustfsadmin \
--print-spark-config
The generated configuration shape is:
spark.sql.catalog.rustfs=org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.rustfs.type=rest
spark.sql.catalog.rustfs.uri=http://127.0.0.1:9000/iceberg
spark.sql.catalog.rustfs.warehouse=rustfs-s3table-smoke
spark.sql.catalog.rustfs.io-impl=org.apache.iceberg.aws.s3.S3FileIO
spark.sql.catalog.rustfs.s3.endpoint=http://127.0.0.1:9000
spark.sql.catalog.rustfs.s3.path-style-access=true
spark.sql.catalog.rustfs.rest.sigv4-enabled=true
spark.sql.catalog.rustfs.rest.signing-name=s3
spark.sql.catalog.rustfs.rest.signing-region=us-east-1
spark.sql.catalog.rustfs.s3.access-key-id=rustfsadmin
spark.sql.catalog.rustfs.s3.secret-access-key=rustfsadmin
Generate the SQL smoke input:
python3 scripts/table-catalog/engine_compatibility.py \
--catalog-name rustfs \
--namespace smoke \
--table events \
--print-spark-sql \
--cleanup
The generated SQL covers namespace creation, table creation, append, refresh, count, and optional cleanup. Until Spark execution is enabled in CI through the explicit live-conformance gate, do not claim Spark support beyond a manually verified run with the exact RustFS build, Spark version, Iceberg version, and expected output recorded.
--print-live-conformance also generates conservative manual probe input for
engines that are not run by default in RustFS CI:
SELECT COUNT(*) command for a
table already created by PyIceberg or Spark. Trino write compatibility is not
claimed.httpfs and iceberg read probe using an operator-supplied
current Iceberg metadata location. The separate duckdb_smoke.py entrypoint
owns the automated generic REST Catalog single-table read/write claim.Generate the full probe set:
python3 scripts/table-catalog/engine_compatibility.py \
--endpoint http://127.0.0.1:9000 \
--warehouse rustfs-s3table-smoke \
--namespace smoke \
--table events \
--metadata-location s3://rustfs-s3table-smoke/tables/table-id/metadata/v1.metadata.json \
--print-live-conformance \
--cleanup
Record the exact engine version, RustFS build, current metadata location, and expected row count when running any of these probes. Treat failures as compatibility findings, not as proof that the server can safely claim broader vendor or engine support.