Skip to content

Commit bd9ec45

Browse files
authored
FEAT: Support for Apache Arrow in Bulk Copy API (#665)
### Work Item / Issue Reference <!-- IMPORTANT: Please follow the PR template guidelines below. For mssql-python maintainers: Insert your ADO Work Item ID below For external contributors: Insert Github Issue number below Only one reference is required - either GitHub issue OR ADO Work Item. --> <!-- mssql-python maintainers: ADO Work Item --> > [AB#46268](https://sqlclientdrivers.visualstudio.com/c6d89619-62de-46a0-8b46-70b92a84d85e/_workitems/edit/46268) <!-- External contributors: GitHub Issue --> > GitHub Issue: #<ISSUE_NUMBER> ------------------------------------------------------------------- ### Summary <!-- Insert your summary of changes below. Minimum 10 characters required. --> This pull request introduces a new high-performance `bulkcopy_arrow` method to the `Cursor` class for bulk loading data directly from Apache Arrow sources, along with several related improvements and refactorings. It also updates type stubs and documentation to reflect the new API, and improves bulk copy authentication handling by refactoring shared logic. **New feature: Arrow-based bulk copy** * Added `Cursor.bulkcopy_arrow(table_name, source)` for efficient bulk loading from Arrow sources (e.g., `pyarrow.Table`, `RecordBatch`, objects exposing the Arrow C Data Interface). This avoids unnecessary Python row materialization and is significantly faster for Arrow-native data. [[1]](diffhunk://#diff-06572a96a58dc510037d5efa622f9bec8519bc1beab13c9f251e97e657a9d4edR12-R23) [[2]](diffhunk://#diff-deceea46ae01082ce8400e14fa02f4b7585afb7b5ed9885338b66494f5f38280R3175-R3338) * The classic `bulkcopy()` method now raises a `TypeError` if given Arrow-shaped data, steering users to the new method. **Codebase refactoring and improvements** * Refactored connection context and Azure AD authentication logic into a shared `_build_pycore_context()` helper, used by both `bulkcopy` and `bulkcopy_arrow`. This ensures consistent authentication handling and reduces code duplication. [[1]](diffhunk://#diff-deceea46ae01082ce8400e14fa02f4b7585afb7b5ed9885338b66494f5f38280R2818-R2942) [[2]](diffhunk://#diff-deceea46ae01082ce8400e14fa02f4b7585afb7b5ed9885338b66494f5f38280L2940-R3076) * Updated type stubs (`mssql_python.pyi`) to include the new `bulkcopy_arrow` method and improved the typing for `bulkcopy`. [[1]](diffhunk://#diff-8b251a7d56f9bb22b686d2b101b0420b1d6cd6934b29edcd408c1f579f7d9e84L7-R7) [[2]](diffhunk://#diff-8b251a7d56f9bb22b686d2b101b0420b1d6cd6934b29edcd408c1f579f7d9e84R212-R241) **Documentation** * Updated the `CHANGELOG.md` to document the new Arrow bulk copy feature and its requirements. **Minor improvements** * Minor code formatting improvement in a test file (`test_bulkcopy_udt_geometry`). <!-- ### PR Title Guide > For feature requests FEAT: (short-description) > For non-feature requests like test case updates, config updates , dependency updates etc CHORE: (short-description) > For Fix requests FIX: (short-description) > For doc update requests DOC: (short-description) > For Formatting, indentation, or styling update STYLE: (short-description) > For Refactor, without any feature changes REFACTOR: (short-description) > For performance improvements PERF: (short-description) > For release related changes, without any feature changes RELEASE: #<RELEASE_VERSION> (short-description) ### Contribution Guidelines External contributors: - Create a GitHub issue first: https://github.com/microsoft/mssql-python/issues/new - Link the GitHub issue in the "GitHub Issue" section above - Follow the PR title format and provide a meaningful summary mssql-python maintainers: - Create an ADO Work Item following internal processes - Link the ADO Work Item in the "ADO Work Item" section above - Follow the PR title format and provide a meaningful summary -->
1 parent 0c230f9 commit bd9ec45

6 files changed

Lines changed: 1652 additions & 147 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,18 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
99
### Added
1010
- New feature: Support for macOS and Linux.
1111
- Documentation: Added API documentation in the Wiki.
12+
- **GH-570:** New `Cursor.bulkcopy_arrow(table_name, source)` method for
13+
high-performance bulk loading directly from Apache Arrow data. Accepts a
14+
`pyarrow.Table`, `RecordBatch`, or `RecordBatchReader`, any object exposing
15+
the Arrow C Data Interface (`__arrow_c_stream__` / `__arrow_c_array__` — e.g.
16+
polars, pandas 2.2+, DuckDB, ADBC results), or an iterable of record batches.
17+
Data is streamed to the server through the Arrow C Data Interface without
18+
materializing intermediate Python row objects, and the GIL is released for
19+
the duration of the network transfer. When the source data already originates
20+
as Arrow, this avoids the Arrow→tuple conversion the classic `bulkcopy()`
21+
path requires (measured ~1.4x–2.7x faster end-to-end for such sources).
22+
`bulkcopy()` now raises `TypeError` steering Arrow inputs to this method.
23+
Requires `mssql-py-core` 0.1.5+.
1224
- Bulk copy now supports `Authentication=ActiveDirectoryServicePrincipal`
1325
via an `entra_id_token_factory` callback registered on the mssql-py-core
1426
connection. The callback is invoked by mssql-tds mid-handshake (FedAuth

0 commit comments

Comments
 (0)