Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLM
Posted by nextime 19 hours ago
Comments
Comment by nextime 5 hours ago
for whoever is reading and rightly say "it seems llm generated", you are right, sorry, wasn't supposed to really post it and it did by error, my bad for not supervision my test while i was doing it.
Comment by SahAssar 18 hours ago
Seems generated. Also why no llamafile?
Comment by nextime 5 hours ago
sorry it was generated and was supposed to be a test and not to be posted. thanks for the llamafile suggestion!
Comment by hypfer 18 hours ago
This feels agentically generated.
The blog, the post here, the (auto?)killed LLM comment.
Comment by nextime 5 hours ago
not the whole blog ( but yes that post ), and yes it WAS autogenerated, it was supposed to be only a test but i did a gating wrong, apologise for it.
Comment by polotics 16 hours ago
what exactly did you mean when xou wrote this paragraph title:
"LiteLLM — a router, not a runtime'
?
Comment by nextime 5 hours ago
I wouldn't wrote as it, sorry, the post was an unintended generated post, wasn't supposed to post for real, until i was going to rewrite it by hand at least.
The meaning for that sentence is that litellm just route requests, it doesn't execute a model.
Comment by anotherCodder 2 hours ago
[flagged]
Comment by nextime 19 hours ago
[flagged]