Embeddings, OCR, reranking and generation don't necessarily run best on the same inference runtime.
SIE uses multiple compute engines, including PyTorch, Flash Attention and SGLang, and selects one per model. The application sees one inference layer. Interested?