Skip to content

PyArrowFile class is not compatible with ABFS uri syntax #2698

Description

@NikitaMatskevich

Apache Iceberg version

0.10.0 (latest release)

Please describe the bug 🐞

Starting from version 20, Pyarrow has support for Azure filesystems.

ABFS URIs have this format: abfs[s]://<file_system>@<account_name>.dfs.core.windows.net//<file_name>

But Pyarrow library expects the following path format for Azure: abfs[s]://<file_system>//<file_name>.

As you see, the part "@<account_name>.<dfs|blob>.core.windows.net" prevents users to use pyarrow file io in Azure environment. This issue CAN be fixed in Pyiceberg by removing account_name part.

The proposed fix is just to start a conversation around the issue. I am not 100% sure how and where this should be fixed.

We know similar issues do not occur with Fsspec file io.

Examples

We have a very basic setup with RestCatalog:

def create_iceberg_catalog():
    CATALOG_URI = "https://lakehouse.../catalog"

    catalog_config = {
        "uri": CATALOG_URI,
        PY_IO_IMPL: "pyiceberg.io.pyarrow.PyArrowFileIO",
        ADLS_ACCOUNT_NAME: "lakehouseaccount",
    }

    return RestCatalog("lakehouse", **catalog_config)

When we create a table "testns.testtable", it is assigned a following location : abfss://lakehouse-azure-bucket@lakehouseaccount.dfs.core.windows.net/testns/testtable

Then, when we try to append data to the table:

data = pa.table(
    {
        "id": pa.array(range(5), type=pa.int32()),  # Ensure 'id' is int32 to match Iceberg schema
        "value": [random.choice(["Heads", "Tails"]) for _ in range(5)],
    }
)
table.append(data)

it throws the following exception:

OSError: ListBlobsByHierarchy failed for prefix='aip_test[/test_table-xxx/metadata/snap-xxx.avro](https://xxx/test_table-xxx.avro)'. GetFileInfo is unable to determine whether the path exists. Azure Error: [InvalidResourceName] 400 The specified resource name contains invalid characters.

This is because exists() method is called:

File [~/.official-venvs/amd64.ipykernel-default.master/lib/python3.12/site-packages/pyiceberg/io/pyarrow.py:368](https://xxx/user/nikita-matckevich/.official-venvs/amd64.ipykernel-default.master/lib/python3.12/site-packages/pyiceberg/io/pyarrow.py#line=367), in PyArrowFile.create(self, overwrite)
    366     if not overwrite and self.exists() is True:

And it expects the uri without "@akehouseaccount.dfs.core.windows.net". When we monkey-patch the PyArrowFile.init everything works fine:

PyArrowFile.old_init = PyArrowFile.__init__
def patched_init(self, location: str, path: str, fs: FileSystem, buffer_size: int = ONE_MEGABYTE):
    # Call the original __init__ method
    self.old_init(location, path, fs, buffer_size)
    self._path = remove_section_between_at_and_slash(path)
    print("Logging: PyArrowFile initialized")
PyArrowFile.__init__ = patched_init

It does not matter how and with which engine the table was created and written before: all pyarrow methods are not working, even those that are on read path, so it will be impossible to scan a non-empty table as well. We tested it by creating a table with fsspec file io and reading it with pyarrow file io.

It is hard to test this behavior with Azurite, because Azurite uris are different and do not contain "@<account_name>" part.

Willingness to contribute

  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time

Activity

  1. kevinjqliu commented on Nov 4, 2025

    @kevinjqliu
    Contributor

    Thanks for opening the issue.

    I wonder if this is an upstream issue. The correct syntax for abfs(s) is

    abfs[s]://<file_system>@<account_name>.dfs.core.windows.net/<path>/<file_name>
    

    according to the docs

    Would be good to check what the expected path is for the pyarrow implementation

  2. kevinjqliu commented on Nov 4, 2025

    @kevinjqliu
    Contributor

    tagging @kyleknap since we were working on fsspec/adlfs together

  3. kevinjqliu commented on Nov 4, 2025

    @kevinjqliu
    Contributor

    Side note, as you suggested, we can try to fix this for our integration. This is the 2nd (3rd?) time where I've seen a FileIO implementation wanting to modify the path uri directly. (HDFS #2291 was the other case I can think of)

  4. NikitaMatskevich commented on Nov 5, 2025

    @NikitaMatskevich
    ContributorAuthor

    I dont know if its was intentional, but right now Pyarrow library expects the following path format for Azure: abfs[s]://<file_system>//<file_name>.

  5. NikitaMatskevich commented on Dec 10, 2025

    @NikitaMatskevich
    ContributorAuthor

    @kyleknap could you look at this issue please? It is blocking some of our clients from using pyiceberg

  6. github-actions commented on Jun 9, 2026

    @github-actions

    This issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.

  7. added this to the PyIceberg 0.12.0 milestone on Jun 29, 2026
  8. krishnakaanchan-png commented on Aug 31, 2026

    @krishnakaanchan-png
    Contributor

    The diagnosis in the issue is not right, and it matters for where the fix goes.

    PyArrow parses the canonical Azure URI fine and takes the account out of it:

    >>> FileSystem.from_uri("abfss://myfs@myacct.dfs.core.windows.net/wh/d.parquet")
    (<pyarrow._azurefs.AzureFileSystem object>, 'myfs/wh/d.parquet')

    So the bug is on our side, in parse_location at pyiceberg/io/pyarrow.py:425:

    return uri.scheme, uri.netloc, f"{uri.netloc}{uri.path}"

    That is written for S3 where the netloc is the bucket. For Azure the netloc is container@account.dfs.core.windows.net, so the account lands inside the path and PyArrow reads the whole first segment as the container name. Container names cannot contain @ or ., so this comes back as a 400 and nothing points at the location.

    There is a second half. _initialize_fs calls self._initialize_azure_fs() with no netloc, unlike the S3 and HDFS ones sitting next to it. So the account can only come from adls.account-name and the one in the location is dropped. That is also why a single FileIO cannot serve two storage accounts.

    FsspecFileIO already does both of these, _ADLS_SCHEMES at fsspec.py:334 and the account taken from the hostname at fsspec.py:276. So this is really about making PyArrowFileIO consistent with it.

    None of the ADLS tests use the account qualified form, they are all abfss://warehouse/<file>, which is why CI never caught this.

    #3884 fixed the path half and is merged now. #3936 does the account half.

    One thing I need your call on. When adls.account-name and the location disagree on the account, should the property win like fsspec does today, or should it raise? Property winning means a catalog spanning two accounts will quietly read from the wrong one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions