Skip to content

fix(output): preserve input index in set_output pandas container - #68

Open
ChrisW09 wants to merge 1 commit into
mainfrom
fix/set-output-preserve-index
Open

ChrisW09 wants to merge 1 commit into
mainfrom
fix/set-output-preserve-index

Conversation

@ChrisW09

@ChrisW09 ChrisW09 commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Closes #60

Problem

With set_output(transform="pandas") (or sklearn.set_config(transform_output="pandas")), Preprocessor.transform returned a DataFrame with a fresh RangeIndex instead of the input's index. Because the method already returns a DataFrame, scikit-learn's set_output wrapper keeps that DataFrame's index rather than restoring the input's. As a result:

  • ColumnTransformer([... Preprocessor ...]).set_output(transform="pandas") raised ValueError: Concatenating DataFrames ... Pandas Indexes that do not match whenever the index was not 0..n-1, e.g. after train_test_split.
  • FeatureUnion([... Preprocessor ...]).set_output(transform="pandas") silently returned 2n rows padded with NaN.

Change

  • pretab/compose/output.py: to_dataframe_output takes a new keyword-only index=None argument, used for the pandas container. Polars has no index and is unchanged.
  • pretab/preprocessor.py: transform passes X.index through.

NumPy and dict inputs still get a default RangeIndex, as before.

Tests

Regression tests added. All four fail on main and pass with this change:

  • tests/integration/test_output_format.py:
    • transform / fit_transform keep a non-default index;
    • ColumnTransformer composition;
    • FeatureUnion composition (n rows, no NaN).
  • tests/compose/test_output.py: unit test for to_dataframe_output(..., index=...).

Full suite: 1442 passed, 59 skipped, 7 xfailed. ruff check and ruff format --check are clean.

Preprocessor.transform built the pandas DataFrame with a fresh RangeIndex.
scikit-learn's set_output wrapper keeps the index of a returned DataFrame,
so the input index was lost: ColumnTransformer(...).set_output("pandas")
raised on a non-default index and FeatureUnion returned 2n NaN-padded rows.
Pass X's index through to_dataframe_output (new keyword-only `index`).

Closes #60

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

set_output(transform="pandas") drops the input index, breaking ColumnTransformer/FeatureUnion composition

1 participant