Skip to main content

Preprocessing

Create and manage data preprocessing pipelines.

All methods are accessed via client.preprocessing.

create_preprocessor()​

POST/v1/preprocessors/create

Create a new preprocessor from a PipelineSpec dict.

Parameters​

namestrRequired
Name of the preprocessor
descriptionstrRequired
Description of the preprocessor
specdictRequired
PipelineSpec dict ({"version": "2.0", "steps": [...]})
sample_dfDataFramedefault: None
Optional sample dataframe for fitting

Returns​

tuple — Tuple of (preprocessor_id, version_id)

Example​

1result = client.preprocessing.create_preprocessor(
2 name="My Resource",
3 description="A description",
4 spec={"version": "2.0", "steps": []},
5 sample_df=df
6)

create_from_pipeline()​

Retired: the platform no longer accepts pipeline binaries. A pickled pipeline runs code wherever it is loaded, so the platform now serves only pipelines it fitted itself from a spec, and answers an uploaded pipeline_binary with HTTP 410 client_pipeline_binaries_retired. This method raises that error locally, before it serialises or sends anything. Call create_preprocessor(name, description, spec, sample_df=df) instead, with the PipelineSpec dict the pipeline was compiled from (\{"version": "2.0", "steps": [...]\}), and the platform fits it. preview_spec dry-runs a spec against a platform dataset before you save it.

Parameters​

namestrRequired
Unused; kept so existing calls get this error, not a TypeError.
descriptionstrRequired
Unused.
pipelineRequired
Unused; it is never serialised or sent.
dfDataFrameRequired
Unused.

Returns​

tuple

Example​

1result = client.preprocessing.create_from_pipeline(
2 name="My Resource",
3 description="A description",
4 pipeline=pipeline,
5 df=df
6)

add_version()​

POST/v1/preprocessors/add

Add a new version to an existing preprocessor.

Parameters​

preprocessor_idstrRequired
ID of the existing preprocessor
specdictRequired
PipelineSpec dict
sample_dfDataFramedefault: None
Optional sample dataframe for fitting
parent_version_idstrdefault: None
Optional parent version for lineage tracking

Returns​

str — The new version_id

Example​

1result = client.preprocessing.add_version(
2 preprocessor_id="pp_abc123",
3 spec={"version": "2.0", "steps": []},
4 sample_df=df,
5 parent_version_id="..."
6)

update_version()​

POST/v1/preprocessors/update-version

Update an existing preprocessor version with a new spec.

Parameters​

version_idstrRequired
ID of the version to update
specdictRequired
Updated PipelineSpec dict
sample_dfDataFramedefault: None
Optional sample dataframe for re-fitting

Returns​

str — The updated version_id

Example​

1result = client.preprocessing.update_version(
2 version_id="version_xyz789",
3 spec={"version": "2.0", "steps": []},
4 sample_df=df
5)

get_version()​

GET/v1/preprocessors/versions/{version_id}

Get metadata for a preprocessor version.

Parameters​

version_idstrRequired
The version ID

Returns​

dict — Version info dict with spec, schemas, etc.

Example​

1result = client.preprocessing.get_version(
2 version_id="version_xyz789"
3)

load_pipeline()​

GET/v1/preprocessors/versions/{version_id}/pipeline

Load a fitted pipeline ready to .transform(). Restricted server-side: the platform serves pipeline binaries to its own train service only, and answers every other caller, a user's API key included, with HTTP 410. For users this method therefore raises until SDK 1.25, which loads a version's fitted state from the platform's /state document instead of a pickle. Until then, to inspect or run a version's transformation: - get_version(version_id) returns its spec and input and output schemas; - preview(version_id, df) runs it on sample data on the platform; - compile_spec(PipelineSpec(**spec)) from xplainable-preprocessing rebuilds it locally from that spec, to .fit() on your own data, when the spec has no custom step. Compiling a custom step runs its code on your machine, so read that code first, or leave the step out. The bytes it receives are unpickled, which can run any code they name, so load only from a platform you trust.

Parameters​

version_idstrRequired
The version ID to load

Returns​

A fitted DataFramePipeline instance

Example​

1result = client.preprocessing.load_pipeline(
2 version_id="version_xyz789"
3)

fit_version()​

POST/v1/preprocessors/versions/{version_id}/fit

Fit a preprocessor version with sample data.

Parameters​

version_idstrRequired
The version ID to fit
dfDataFrameRequired
Dataframe to fit on

Returns​

dict — Fit result dict with schemas and status

Example​

1result = client.preprocessing.fit_version(
2 version_id="version_xyz789",
3 df=df
4)

preview()​

POST/v1/preprocessors/versions/{version_id}/preview

Preview pipeline transformation on sample data.

Parameters​

version_idstrRequired
The version ID to preview
dfDataFrameRequired
Sample dataframe

Returns​

dict — Preview dict with deltas, schemas, and samples

Example​

1result = client.preprocessing.preview(
2 version_id="version_xyz789",
3 df=df
4)

list_preprocessors()​

GET/v1/preprocessors/teams/{team_id}

List all preprocessors for a team.

Parameters​

team_idstrdefault: None
Optional team ID (uses session team_id if not provided)

Returns​

list — List of preprocessor information

Example​

1result = client.preprocessing.list_preprocessors(
2 team_id="team_abc123"
3)

get_preprocessor()​

GET/v1/preprocessors/{preprocessor_id}

Get detailed information about a preprocessor.

Parameters​

preprocessor_idstrRequired
ID of the preprocessor

Returns​

PreprocessorInfo — Preprocessor information

Example​

1result = client.preprocessing.get_preprocessor(
2 preprocessor_id="pp_abc123"
3)

check_signature()​

POST/v1/preprocessors/check-signature

Check whether a preprocessor version's output columns match a model version's training columns (ignoring id and target columns). Use before link_preprocessor or deploying a model with a preprocessor attached: a mismatch means inference rows would be transformed into a schema the model was not trained on.

Parameters​

preprocessor_version_idstrRequired
The preprocessor version whose output schema to check.
model_version_idstrRequired
The model version whose training columns are the expected schema.

Returns​

dict — Dict with signatures_match (bool).

Example​

1result = client.preprocessing.check_signature(
2 preprocessor_version_id="ppv_xyz789",
3 model_version_id="version_xyz789"
4)

delete_version()​

DELETE/v1/preprocessors/versions/{version_id}

Delete a preprocessor version.

Parameters​

version_idstrRequired
The version ID to delete

Returns​

dict — Deletion result dict

Example​

1result = client.preprocessing.delete_version(
2 version_id="version_xyz789"
3)

create_preprocessor_from_spec()​

Create a new preprocessor from a PipelineSpec dict. The spec should follow the PipelineSpec format: {"version": "2.0", "steps": [{"id": "...", "type": "...", "columns": [...], "params": {...}}]} Use preprocessing_list_available_transformers to see available transformer types and their parameters.

Parameters​

namestrRequired
Name of the preprocessor
descriptionstrRequired
Description of the preprocessor
specdictRequired
PipelineSpec dict
sample_datalistdefault: None
Optional sample data as a list of row dicts (JSON records)

Returns​

dict — Dict with preprocessor_id and version_id

Example​

1result = client.preprocessing.create_preprocessor_from_spec(
2 name="My Resource",
3 description="A description",
4 spec={"version": "2.0", "steps": []},
5 sample_data=[{"col1": 1, "col2": "a"}]
6)

add_version_from_spec()​

Add a new version to an existing preprocessor.

Parameters​

preprocessor_idstrRequired
ID of the existing preprocessor
specdictRequired
PipelineSpec dict
sample_datalistdefault: None
Optional sample data as a list of row dicts (JSON records)
parent_version_idstrdefault: None
Optional parent version for lineage tracking

Returns​

dict — Dict with version_id

Example​

1result = client.preprocessing.add_version_from_spec(
2 preprocessor_id="pp_abc123",
3 spec={"version": "2.0", "steps": []},
4 sample_data=[{"col1": 1, "col2": "a"}],
5 parent_version_id="..."
6)

update_version_from_spec()​

Update an existing preprocessor version with a new spec.

Parameters​

version_idstrRequired
ID of the version to update
specdictRequired
Updated PipelineSpec dict
sample_datalistdefault: None
Optional sample data as a list of row dicts (JSON records)

Returns​

dict — Dict with version_id

Example​

1result = client.preprocessing.update_version_from_spec(
2 version_id="version_xyz789",
3 spec={"version": "2.0", "steps": []},
4 sample_data=[{"col1": 1, "col2": "a"}]
5)

fit_version_from_data()​

Fit a preprocessor version with sample data.

Parameters​

version_idstrRequired
The version ID to fit
sample_datalistRequired
Sample data as a list of row dicts (JSON records)

Returns​

dict — Fit result dict with schemas and status

Example​

1result = client.preprocessing.fit_version_from_data(
2 version_id="version_xyz789",
3 sample_data=[{"col1": 1, "col2": "a"}]
4)

preview_from_data()​

Preview pipeline transformation on sample data.

Parameters​

version_idstrRequired
The version ID to preview
sample_datalistRequired
Sample data as a list of row dicts (JSON records)

Returns​

dict — Preview dict with deltas, schemas, and samples

Example​

1result = client.preprocessing.preview_from_data(
2 version_id="version_xyz789",
3 sample_data=[{"col1": 1, "col2": "a"}]
4)

preview_spec()​

POST/v1/preprocessors/preview-spec

Dry-run a preprocessing spec against a platform dataset without persisting anything. Each step is applied in order on a sample of the dataset and reported with its column delta (dropped/added/updated), before/after statistics and warnings. The same safety guard the platform enforces on create/add runs here first, so a spec that would be rejected (row-collapsing steps, dropping or mutating the target) is flagged in safety_errors before you try to save it. Use this to iterate on a spec, then persist it with create_preprocessor_from_spec / add_version_from_spec.

Parameters​

dataset_idstrRequired
Platform dataset to preview against (see datasets_list_team_datasets)
specdictRequired
PipelineSpec dict: {"version": "2.0", "steps": [{"id", "type", "columns", "params", "description"}]}
target_columnstrdefault: None
Optional prediction target; enables the drop/mutate-label checks
max_rowsintdefault: 10000
Rows to sample for the preview (default 10000)

Returns​

dict — Dict with steps (per-step delta/stats/warnings/error), safety_errors, rows_before, rows_after and output_columns

Example​

1result = client.preprocessing.preview_spec(
2 dataset_id="ds_abc123",
3 spec={"version": "2.0", "steps": []},
4 target_column="target",
5 max_rows=1
6)

list_available_transformers()​

List all available preprocessing transformers with their parameters. Returns a catalog of transformer types that can be used in PipelineSpec steps, including their constructor parameters and descriptions. Custom-code steps (type "custom") are not offered: the platform refuses a spec with custom code it has not reviewed (422 custom_steps_disabled), so build every step from the types listed.

Returns​

str — Formatted string describing all available transformers

Example​

1result = client.preprocessing.list_available_transformers()

delete_preprocessor()​

DELETE/v1/preprocessors/{preprocessor_id}

Delete a preprocessor and all its versions.

Parameters​

preprocessor_idstrRequired
The preprocessor ID to delete

Returns​

dict — Deletion result dict

Example​

1result = client.preprocessing.delete_preprocessor(
2 preprocessor_id="pp_abc123"
3)