Preprocessing
Create and manage data preprocessing pipelines.
All methods are accessed via client.preprocessing.
create_preprocessor()
Create a new preprocessor from a PipelineSpec dict.
Parameters
NoneReturns
tuple — Tuple of (preprocessor_id, version_id)
Example
create_from_pipeline()
Retired: the platform no longer accepts pipeline binaries. A pickled pipeline runs code wherever it is loaded, so the platform now serves only pipelines it fitted itself from a spec, and answers an uploaded pipeline_binary with HTTP 410 client_pipeline_binaries_retired. This method raises that error locally, before it serialises or sends anything. Call create_preprocessor(name, description, spec, sample_df=df) instead, with the PipelineSpec dict the pipeline was compiled from (\{"version": "2.0", "steps": [...]\}), and the platform fits it. preview_spec dry-runs a spec against a platform dataset before you save it.
Parameters
Returns
tuple
Example
add_version()
Add a new version to an existing preprocessor.
Parameters
NoneNoneReturns
str — The new version_id
Example
update_version()
Update an existing preprocessor version with a new spec.
Parameters
NoneReturns
str — The updated version_id
Example
get_version()
Get metadata for a preprocessor version.
Parameters
Returns
dict — Version info dict with spec, schemas, etc.
Example
load_pipeline()
Load a fitted pipeline ready to .transform(). Restricted server-side: the platform serves pipeline binaries to its own train service only, and answers every other caller, a user's API key included, with HTTP 410. For users this method therefore raises until SDK 1.25, which loads a version's fitted state from the platform's /state document instead of a pickle. Until then, to inspect or run a version's transformation: - get_version(version_id) returns its spec and input and output schemas; - preview(version_id, df) runs it on sample data on the platform; - compile_spec(PipelineSpec(**spec)) from xplainable-preprocessing rebuilds it locally from that spec, to .fit() on your own data, when the spec has no custom step. Compiling a custom step runs its code on your machine, so read that code first, or leave the step out. The bytes it receives are unpickled, which can run any code they name, so load only from a platform you trust.
Parameters
Returns
A fitted DataFramePipeline instance
Example
fit_version()
Fit a preprocessor version with sample data.
Parameters
Returns
dict — Fit result dict with schemas and status
Example
preview()
Preview pipeline transformation on sample data.
Parameters
Returns
dict — Preview dict with deltas, schemas, and samples
Example
list_preprocessors()
List all preprocessors for a team.
Parameters
NoneReturns
list — List of preprocessor information
Example
get_preprocessor()
Get detailed information about a preprocessor.
Parameters
Returns
PreprocessorInfo — Preprocessor information
Example
check_signature()
Check whether a preprocessor version's output columns match a model version's training columns (ignoring id and target columns). Use before link_preprocessor or deploying a model with a preprocessor attached: a mismatch means inference rows would be transformed into a schema the model was not trained on.
Parameters
Returns
dict — Dict with signatures_match (bool).
Example
delete_version()
Delete a preprocessor version.
Parameters
Returns
dict — Deletion result dict
Example
create_preprocessor_from_spec()
Create a new preprocessor from a PipelineSpec dict. The spec should follow the PipelineSpec format: {"version": "2.0", "steps": [{"id": "...", "type": "...", "columns": [...], "params": {...}}]} Use preprocessing_list_available_transformers to see available transformer types and their parameters.
Parameters
NoneReturns
dict — Dict with preprocessor_id and version_id
Example
add_version_from_spec()
Add a new version to an existing preprocessor.
Parameters
NoneNoneReturns
dict — Dict with version_id
Example
update_version_from_spec()
Update an existing preprocessor version with a new spec.
Parameters
NoneReturns
dict — Dict with version_id
Example
fit_version_from_data()
Fit a preprocessor version with sample data.
Parameters
Returns
dict — Fit result dict with schemas and status
Example
preview_from_data()
Preview pipeline transformation on sample data.
Parameters
Returns
dict — Preview dict with deltas, schemas, and samples
Example
preview_spec()
Dry-run a preprocessing spec against a platform dataset without persisting anything. Each step is applied in order on a sample of the dataset and reported with its column delta (dropped/added/updated), before/after statistics and warnings. The same safety guard the platform enforces on create/add runs here first, so a spec that would be rejected (row-collapsing steps, dropping or mutating the target) is flagged in safety_errors before you try to save it. Use this to iterate on a spec, then persist it with create_preprocessor_from_spec / add_version_from_spec.
Parameters
None10000Returns
dict — Dict with steps (per-step delta/stats/warnings/error), safety_errors, rows_before, rows_after and output_columns
Example
list_available_transformers()
List all available preprocessing transformers with their parameters. Returns a catalog of transformer types that can be used in PipelineSpec steps, including their constructor parameters and descriptions. Custom-code steps (type "custom") are not offered: the platform refuses a spec with custom code it has not reviewed (422 custom_steps_disabled), so build every step from the types listed.
Returns
str — Formatted string describing all available transformers
Example
delete_preprocessor()
Delete a preprocessor and all its versions.
Parameters
Returns
dict — Deletion result dict