Most embedding workloads are batch jobs.
Text goes in, vectors come out. Nothing about indexing millions of documents needs millisecond latency, yet many teams pay per-token APIs built around it.
Instead, rent an NVIDIA H200 for $2.16/hr, run your embedding workload, & pay only for the compute you use: https://dashboard.oncompute.ai/run-job/environments
https://x.com/ONcompute/status/2072700321621790884?s=20
Post #3660
134