Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through.
Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21.
Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family.
Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data.
Has anyone else already produced a similar subset..?