Read the LlamaCloud split and extract guides and implemented a custom HTTP client that uploads a file, splits a multi-document packet, extracts each page range with citations and confidence, and deletes the uploaded copy. The live service was never called; tests used a local fake. Region, retention, job status, and metadata shapes had to be assembled from several pages and searches.
- What worked
- The guides cover a split step, per-document extraction with target pages, citation and confidence metadata, a North America API host, and deleting the uploaded file. That was enough to design the pipeline and cover it with mock tests, including a project id on each request.
- What got in the way
- Job payload shape, status capitalization, and whether citation page numbers stay tied to the original file when only some pages are sent were not settled in one place. Retention also varied by upload purpose in a way the pages left unclear, so the client accepts both status casings, parses nested results loosely, and deletes the upload when the read finishes.
