Repository navigation
PyArrowFile class is not compatible with ABFS uri syntax #2698
Description
Activity
Thanks for opening the issue.
I wonder if this is an upstream issue. The correct syntax for abfs(s) is
abfs[s]://<file_system>@<account_name>.dfs.core.windows.net/<path>/<file_name>according to the docs
Would be good to check what the expected path is for the pyarrow implementation
tagging @kyleknap since we were working on fsspec/adlfs together
Side note, as you suggested, we can try to fix this for our integration. This is the 2nd (3rd?) time where I've seen a FileIO implementation wanting to modify the path uri directly. (HDFS #2291 was the other case I can think of)
NikitaMatskevich commented
on Nov 5, 2025 ContributorAuthorMore actionsI dont know if its was intentional, but right now Pyarrow library expects the following path format for Azure: abfs[s]://<file_system>//<file_name>.
NikitaMatskevich commented
on Dec 10, 2025 ContributorAuthorMore actions@kyleknap could you look at this issue please? It is blocking some of our clients from using pyiceberg
This issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.
krishnakaanchan-png commented
on Aug 31, 2026 ContributorMore actionsThe diagnosis in the issue is not right, and it matters for where the fix goes.
PyArrow parses the canonical Azure URI fine and takes the account out of it:
>>> FileSystem.from_uri("abfss://myfs@myacct.dfs.core.windows.net/wh/d.parquet") (<pyarrow._azurefs.AzureFileSystem object>, 'myfs/wh/d.parquet')
So the bug is on our side, in
parse_locationatpyiceberg/io/pyarrow.py:425:return uri.scheme, uri.netloc, f"{uri.netloc}{uri.path}"
That is written for S3 where the netloc is the bucket. For Azure the netloc is
container@account.dfs.core.windows.net, so the account lands inside the path and PyArrow reads the whole first segment as the container name. Container names cannot contain@or., so this comes back as a 400 and nothing points at the location.There is a second half.
_initialize_fscallsself._initialize_azure_fs()with no netloc, unlike the S3 and HDFS ones sitting next to it. So the account can only come fromadls.account-nameand the one in the location is dropped. That is also why a single FileIO cannot serve two storage accounts.FsspecFileIOalready does both of these,_ADLS_SCHEMESatfsspec.py:334and the account taken from the hostname atfsspec.py:276. So this is really about makingPyArrowFileIOconsistent with it.None of the ADLS tests use the account qualified form, they are all
abfss://warehouse/<file>, which is why CI never caught this.#3884 fixed the path half and is merged now. #3936 does the account half.
One thing I need your call on. When
adls.account-nameand the location disagree on the account, should the property win like fsspec does today, or should it raise? Property winning means a catalog spanning two accounts will quietly read from the wrong one.
Apache Iceberg version
0.10.0 (latest release)
Please describe the bug 🐞
Starting from version 20, Pyarrow has support for Azure filesystems.
ABFS URIs have this format: abfs[s]://<file_system>@<account_name>.dfs.core.windows.net//<file_name>
But Pyarrow library expects the following path format for Azure: abfs[s]://<file_system>//<file_name>.
As you see, the part "@<account_name>.<dfs|blob>.core.windows.net" prevents users to use pyarrow file io in Azure environment. This issue CAN be fixed in Pyiceberg by removing account_name part.
The proposed fix is just to start a conversation around the issue. I am not 100% sure how and where this should be fixed.
We know similar issues do not occur with Fsspec file io.
Examples
We have a very basic setup with RestCatalog:
When we create a table "testns.testtable", it is assigned a following location : abfss://lakehouse-azure-bucket@lakehouseaccount.dfs.core.windows.net/testns/testtable
Then, when we try to append data to the table:
it throws the following exception:
This is because exists() method is called:
And it expects the uri without "@akehouseaccount.dfs.core.windows.net". When we monkey-patch the PyArrowFile.init everything works fine:
It does not matter how and with which engine the table was created and written before: all pyarrow methods are not working, even those that are on read path, so it will be impossible to scan a non-empty table as well. We tested it by creating a table with fsspec file io and reading it with pyarrow file io.
It is hard to test this behavior with Azurite, because Azurite uris are different and do not contain "@<account_name>" part.
Willingness to contribute