I like the unglamorous half of AI work: the ingestion, the queues, the retries. Most of what I build ends in a model call, but the part that decides whether it works is everything before it: backfilling 280k procurement records through a rate limited job queue, checkpointing extracted document text so a failed embedding run resumes instead of starting over, grouping raw solar telemetry into events an operator can actually act on. If it cannot survive being run unattended at 3am, it is a demo, not a system.
The work I am proudest of is a number I made worse. A classifier of mine hit 0.91 ROC-AUC, which felt great until I traced it to study overlap across splits and test set exposure during feature selection. Rebuilt with cohort grouped cross validation and fold specific selection, it reported 0.64. That is the honest number, and finding it taught me more than the first one would have. I would rather ship something defensible than something impressive.
The Mandalorian thing is about craft, not cosplay. A covert of people who are very good at one narrow trade, who maintain their own gear, who take a creed seriously enough to be inconvenienced by it. Beskar gets reforged, never thrown out. That is a fair description of how I feel about a codebase you intend to keep.