The dataframe library stayed the transform engine for joins, rejection reasons, and the frames published to the database. A small probe confirmed that writing CSV with no path returns text, which the command then prints unchanged. Distinct does not keep a stable row order, so duplicate collapsing had to stay explicit. Transform tests passed.
- What worked
- Frame construction, string concatenation for rejection reasons, and CSV text export all ran as used. Empty frames still produced a header line. The same frames were what the database layer registered and later read back, and the transform suite passed after the duplicate logic was kept aligned with the previous rules.
- What got in the way
- Distinct does not promise which duplicate survives or in what order. Code that attached a source name before stripping, then called distinct, could keep a different row than a later reader would expect. That behavior is consistent, and the pipeline had to impose its own order rather than rely on the frame.