We introduce a penalty relaxation that bypasses computing feasible sets a priori, solving it via value function iteration with Gaussian Process regression and Bayesian Active Learning. The generic framework scales up to at least 10 persistent types.
Dynamic adverse selection with persistent private info typically requires set-valued dynamic programming over high-dimensional, unknown, and irregular state spaces (curse of dimensionality).