A great many problems presented as needing deep learning are forecasting, classification, or ranking problems that a well-specified classical model handles better. Better, not merely cheaper: simpler models are easier to explain to the people who must act on their output, easier to debug when they drift, and far easier to retrain on a schedule someone can actually maintain.
The decisive work is usually upstream of the model. Where the data comes from, how late it arrives, what happens when a source changes shape, whether the features available at training time will genuinely be available at prediction time — these determine the ceiling on performance long before the choice of algorithm does. A model trained on information it will not have in production is a common and expensive failure, and it looks excellent in evaluation.
So the approach is unglamorous on purpose: establish a baseline that anyone can reason about, make the pipeline that feeds it reliable and repeatable, and only add model complexity where it earns a measurable improvement against a decision someone is actually making.