Did this on Langfuse’s CI recently. Told Fable to make our pipelines 30s faster end-to-end, iterating on real CI timings with old runs as the baseline.
10% speed up over night - quite the win for an already very fast pipeline (5min).
I'm having a lot of success giving GPT more ambitious /goal criteria, even if I don't expect the model to ever hit the goal.
E.g., "Improve performance by 10%" rather than 1%. Tell it to be ambitious, think big, etc.
Then put up PRs as it goes and merge discrete improvements.