Support LocationProviders like the Java Iceberg Reference Implementaiton #861

corleyma · 2024-06-26T18:21:18Z

Feature Request / Improvement

The Iceberg Java API supports customizing where data files are written using LocationProviders. By default, Iceberg writes data files following a Hive schema. However, table properties can be used to switch to using a built-in ObjectStoreLocationProvider that is optimized to maximize throughput on object storage. It's also possible to provide your own custom LocationProviders.

I think for write parity between pyiceberg and Spark, it's important for pyiceberg to at least support the ObjectStoreLocationProvider functionality. Here's the Java implementation, and here's the current logic in pyiceberg which seems to only mimic the default locationprovider.

While AWS S3 has since revised their guidance about the necessity of injecting randomness into prefixes, that guidance still holds for other object stores like GCS, so in addition to being required for parity I think this feature is still important for performance too.

kevinjqliu · 2024-06-27T00:41:13Z

+1 on adding LocationProviders. It's useful for configuring writes for data files and gives flexibility in writing.
Would be great to make it pluggable just like FileIO

smaheshwar-pltr · 2024-12-19T03:29:23Z

I think this is a great idea! Would love to work on this if no one is assigned yet 😄

Fokko · 2024-12-19T09:36:38Z

@smaheshwar-pltr I don't think anyone started on this, feel free to pick this up, I'm happy to review!

smaheshwar-pltr · 2024-12-23T15:32:03Z

Great! I've put up #1452 that should address this

corleyma mentioned this issue Oct 28, 2024

when using SQL catalog, table location is not optional #1254

Closed

smaheshwar-pltr linked a pull request Dec 20, 2024 that will close this issue

Support Location Providers #1452

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Support LocationProviders like the Java Iceberg Reference Implementaiton #861

Support LocationProviders like the Java Iceberg Reference Implementaiton #861

corleyma commented Jun 26, 2024

kevinjqliu commented Jun 27, 2024

smaheshwar-pltr commented Dec 19, 2024

Fokko commented Dec 19, 2024

smaheshwar-pltr commented Dec 23, 2024

Support LocationProviders like the Java Iceberg Reference Implementaiton #861

Support LocationProviders like the Java Iceberg Reference Implementaiton #861

Comments

corleyma commented Jun 26, 2024

Feature Request / Improvement

kevinjqliu commented Jun 27, 2024

smaheshwar-pltr commented Dec 19, 2024

Fokko commented Dec 19, 2024

smaheshwar-pltr commented Dec 23, 2024