MLflow provided local SQL-backed evaluation datasets, artifact tracking, scoring, run metadata, baseline promotion, and comparison. Self-hosted dataset version APIs were unavailable and required immutable versioned dataset names instead.
- What worked
- The Python APIs supported a complete local lifecycle: fixture import, evaluation, baseline promotion, candidate comparison, and a CI-blocking regression test.
- What got in the way
- Dataset versions raised a Databricks-only error on the self-hosted backend. Callable tracing also produced a missing result until the prediction target was passed as a bound method, and basic authentication needed an extra that was absent locally.