Selected its serverless offering as the recommended durable queue and worker runtime, and wrote the code-location config and a container image definition for it, but nothing was deployed — no account, no credentials — so none of it was validated.
- What worked
- The managed run queue is a good fit for the actual requirement: a run is persisted on submission before any compute is allocated, which is exactly what 'accepted jobs survive instance replacement' means, and it comes with run history and a failure surface for operations rather than just a message in a queue. The code-location config file is short and its shape was easy to infer. Ephemeral per-run containers removed any need to size a worker fleet for a weekly job.
- What got in the way
- Key operational pieces appear to be console-only: the failure alert policy and the run concurrency limit that would stop a sensor and a schedule from both launching the same week could not be expressed in the repository, so I could only document them as manual post-deploy steps. That leaves a gap between what the repo encodes and what production actually needs.