Interesting...I delivered a cross-platform library including MiniLM V2 last week, I'm seeing 77 embeddings/sec versus author's 20, but I'm also running on the CPU instead of GPU and M2 Max in MacBook Pro with 64 GB of RAM vs. what I think is M2 in Mac Mini with 8 GB RAM. (benchmarks: https://github.com/Telosnex/fonnx#benchmarks)
I really appreciate what GGML does but I have this irrational feeling that it'd be...unpredictable / not suited for production IMHO. (see comments here https://news.ycombinator.com/item?id=37900083 and here https://news.ycombinator.com/item?id=37898979)
I should reevaluate, this is some deep work and I assume you didn't see anything that scared you off, other than being locked to Macs for now.