Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLM

Posted by nextime 19 hours ago

Counter12Comment7OpenOriginal

Comments

Comment by nextime 5 hours ago

for whoever is reading and rightly say "it seems llm generated", you are right, sorry, wasn't supposed to really post it and it did by error, my bad for not supervision my test while i was doing it.

Comment by SahAssar 18 hours ago

Seems generated. Also why no llamafile?

Comment by nextime 5 hours ago

sorry it was generated and was supposed to be a test and not to be posted. thanks for the llamafile suggestion!

Comment by hypfer 18 hours ago

This feels agentically generated.

The blog, the post here, the (auto?)killed LLM comment.

Comment by nextime 5 hours ago

not the whole blog ( but yes that post ), and yes it WAS autogenerated, it was supposed to be only a test but i did a gating wrong, apologise for it.

Comment by polotics 16 hours ago

what exactly did you mean when xou wrote this paragraph title: "LiteLLM — a router, not a runtime' ?

Comment by nextime 5 hours ago

I wouldn't wrote as it, sorry, the post was an unintended generated post, wasn't supposed to post for real, until i was going to rewrite it by hand at least.

The meaning for that sentence is that litellm just route requests, it doesn't execute a model.

Comment by anotherCodder 2 hours ago

[flagged]

Comment by nextime 19 hours ago

[flagged]