Can gzip be a language model?
Posted by networked 17 hours ago
Comments
Comment by jll29 14 hours ago
gzip -9 sports.txt testfile.txt
gzip -9 politics.txt testfile.txt
gzip -9 business.txt testfile.txt
(ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)The test file belongs to the topic with the smallest size *.gz file.
Witten's group at Waikato uni were perhaps the first to work on this.
Also check out the Hutter prize if you are interested in this.
Comment by stingraycharles 12 hours ago
I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages.
I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language.
Absolutely totally not the best approach, but very fast and super simple to implement.
Comment by ape4 9 hours ago
Comment by wongarsu 8 hours ago
Also some languages have a lot of prefixes and suffixes on their verbs or even nouns, which dilutes your list of 1000 words by just adding the same common words over and over again with different suffixes designating grammatical tense, grammatical gender, etc.
The gzip version sounds more general and more obviously correct
Comment by basilgohar 8 hours ago
Comment by ashkankiani 6 hours ago
Comment by dspillett 5 hours ago
Comment by giancarlostoro 4 hours ago
Comment by BobaFloutist 5 hours ago
Comment by colejohnson66 4 hours ago
Comment by thesz 7 hours ago
[1] https://en.wikipedia.org/wiki/Byte-pair_encoding
It naturally takes care of common prefixes and suffixes.
It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.
Comment by kragen 8 hours ago
Comment by wodenokoto 3 hours ago
Comment by ignoramous 1 hour ago
FSST is based on a fixed size (255 items) dictionary of high frequency variable length strings/substrings (learned from the corpus) encoded as one byte.
Comment by actionfromafar 9 hours ago
Comment by LPisGood 12 hours ago
Also, I’ve never seen “ass.” Used to shorten “aside” — I typically use N.B. but perhaps only for important ones.
Comment by shoo 12 hours ago
Comment by matzf 11 hours ago
Comment by chrisweekly 10 hours ago
Comment by arrowsmith 12 hours ago
Comment by woadwarrior01 12 hours ago
https://en.wikipedia.org/wiki/Normalized_compression_distanc...
Comment by chris_va 9 hours ago
https://en.wikipedia.org/wiki/Baconian_theory_of_Shakespeare...
By looking at mutual information from different authors on the same topic vs same author on different topics. As I recall, it convincingly disproved the hypothesis.
Comment by mrtnmcc 3 hours ago
Comment by myrmidon 9 hours ago
Really interesting approach though.
Comment by Lerc 9 hours ago
Comment by ape4 9 hours ago
Comment by jjtheblunt 3 hours ago
Comment by anthk 3 hours ago
Comment by m-hodges 11 hours ago
Comment by Culonavirus 15 hours ago
Comment by wolfi1 15 hours ago
Comment by shezi 14 hours ago
Looks pretty profitable to me.
Comment by amiga386 14 hours ago
That said, Windows users should use 7-Zip. Better compression format, unpacks more kinds of archives
Comment by TonyTrapp 9 hours ago
Comment by xxs 14 hours ago
Please no - no native zstd support. NanaZip is the better option (it's a different build of 7-zip) and it's available at windows store.
> Better compression format, unpacks more kinds of archives
winrar has supported zstd for 5 years[0]
In short - Everyone should be using zstd, and 7-zip does not support it.
[0]: https://www.win-rar.com/singlenewsview.html?&L=0&tx_ttnews%5...
Comment by tnelsond4 12 hours ago
Comment by cgio 12 hours ago
|gap |gzip |bz2 |lzma | |---------|----------|-----|--------| |0 |2.7% |18.9%|*0.9%*| |8 KB |2.4% |17.8%|0.9% | |*40 KB*|*94.4%* |17.7%|0.4% | |1 MB |*104.3%*|19.7%|*0.9%*|
Comment by notpushkin 10 hours ago
I don’t see zstd in your comparison?
Comment by Sweepi 9 hours ago
|gap |gzip |bz2 |lzma |
|---------|----------|-----|--------|
|0 |2.7% |18.9%| *0.9%*|
|8 KB |2.4% |17.8%| 0.9% |
|*40 KB* |*94.4%* |17.7%| 0.4% |
|1 MB |*104.3%* |19.7%| *0.9%*|Comment by ndriscoll 9 hours ago
Comment by pixl97 8 hours ago
7z is now built into W11 right click so that or zip is what will be used by default anyway.
Comment by hypercube33 36 minutes ago
Comment by aleph_minus_one 13 hours ago
Why?
Comment by shawabawa3 13 hours ago
Comment by xxs 13 hours ago
On a more realistic note: few years back, I've added zstd compression to our log subsystem (hand written direct buffers, native code, in-process, java). For the same CPU utilization if provides twice dense compression compared to regular [-6] gzip (the topic in the title). Zstd is =much= faster on decompression as well, and it this case - unparalleledly better as it uses twice less disk.
zstd is 'silicon valley' (the tv show) - life imitates fiction, except entirely open source
Comment by gsich 6 hours ago
Comment by dd8601fn 14 hours ago
Comment by cavoirom 10 hours ago
Comment by zamadatix 5 hours ago
E.g. when SpaceX started designing the Falcon 9 the losses were astronomical (pun intended) but a car dealership running more of a profit at the same time doesn't have much to say about which business is doing better.
Comment by make3 5 hours ago
Comment by zamadatix 5 hours ago
Who knows though - maybe I missed that German company financials are also funny :p. Wouldn't be the first time something flew well over my head.
Comment by jurgenburgen 14 hours ago
Comment by Betelbuddy 13 hours ago
Comment by firtoz 13 hours ago
Comment by tecleandor 13 hours ago
The company doing the software distribution, is located in Berlin. The Managing Directors for that company seem to have Turkish names, but I don't know if they're Turkish.
BTW, Looking for some info I just found a website [0], clearly AI generated (but not necessarily meaning the content is false) claiming Eugene Roshal had severe kidney failure this past month, and he's waiting for surgery. They're asking for donations. There are some names on who's theoretically behind it [1] but they don't link to any LinkedIn profile or personal site. I can't find any other references. The BTC wallet they're using for donations hasn't seen any traffic ever. BE WARY, SMELLS FISHY.
--
0: https://eugeneroshal.org/
1: https://eugeneroshal.org/about/Comment by pshan 8 hours ago
Comment by vova_hn2 8 hours ago
I suspect that someone's running a script that looks for public (-ish, Eugene has a Wikipedia page at least) figures without active social media presence and creates fake donation websites with AI-generated texts and pictures.
If my suspicion is correct, this is one of the most evil scams I can imagine.
Comment by grezql 9 hours ago
edit: can you write in brackets after the URL that it may be scam. I think alot of people may just go to the URL without reading your warning
Comment by GodelNumbering 14 hours ago
Comment by kevinrineer 4 hours ago
Comment by mg 15 hours ago
give it a normal text prompt, and it
continues that prompt by searching
for the byte sequences that compress
best.
One moment, how are we supposed to know how well that search was done? There is no way to search a meaningful part of the search space.So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.
The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?
Comment by shoo 14 hours ago
It's unclear if this is very useful.
The reason it may not be very useful is that one of Deflate's ingredients is a pass that replaces repeated substrings with backreferences to the earlier occurrence in the plaintext input stream.
E.g. suppose we want to find an n=200 byte sequence x that minimises len(gzip(context+prompt+x)).
If there exists any 200 byte sequence y such that prompt+y is a substring of context, then Deflate can encode prompt+y as a backreference to that earlier sequence - it needs to store a match-length & a distance-length, encoded using its Huffman trees. This candidate solution y may not be a global minima to our stated objective function, but if not, it's probably going to be a very good near-optimal approximate solution.
Taking a step back, repeating huge chunks of the input context produces something that's great for minimising compressed output size but doesn't seem particularly helpful as a generative model.
edit:
Yep, I tried it out by running an experiment. Searching for the prompt in the context & then copying the following text as the solution produces solutions that are much better, in the sense of minimising the compressed output length, than beam search, while also being unhelpful as a generative tool.
With the same example as the blog post:
context: first 30,000 bytes of tinyshakespeare.txt
prompt: 'MENENIUS:\n'
Let x denote a solution, x is a string of length 200.Let L(x) denote len(gzip(context+prompt+x)), our objective function
Let's call the proposed search method of searching for the prompt in the input rfind (after python's str.rfind).
Then we have
search method soln soln length feasible? objective value search time (wall clock, s)
------------- ---- ----------- --------- --------------- ---------------------------
emptystring "" 0 no 13,023 0.04s
gzipt beam search see blog post 200 yes 13,051 11.93s
rfind see below 200 yes 13,026 0.04s
So 'rfind' is finding a solution that does a better job of minimising the objective function -- it only takes 3 bytes more to encode than the infeasible emptystring solution, and costs 25 fewer bytes than the solution found by the beam search implemented by gzipt per the blog post.Here's the solution 'generated' by rfind copying and pasting from the input context, starting from the rightmost occurrence of "MENENIUS:"
MENENIUS:
O, true-bred!
First Senator:
Your company to the Capitol; where, I know,
Our greatest friends attend us.
TITUS:
COMINIUS:
Noble Marcius!
First Senator:
MARCIUS:
Nay, let them follow:
The Volsces
Here's the code for 'rfind' - our complete 'generative algorithm': def find_candidate_solution_from_context(context, prompt, length):
n = len(context)
i = context.rfind(prompt, 0, n-length)
if i < 0:
return b''
i += len(prompt)
return context[i:i+length]
Can hook it into gzipt.py by adding this line after out is defined, but before the beam search begins out += find_candidate_solution_from_context(corpus_window, prompt, length)Comment by StilesCrisis 11 hours ago
Basically I think the entire premise falls apart due to that choice--they forced an interesting-looking outcome by adjusting the algorithm until gzip started picking random slabs of letters instead of ever-larger repeating runs.
Comment by Ohentis 8 hours ago
Comment by im_down_w_otp 9 hours ago
Comment by StilesCrisis 8 hours ago
Comment by adamgordonbell 3 hours ago
https://bellard.org/ts_zip/ https://corecursive.com/the-hutter-prize/
Comment by montebicyclelo 14 hours ago
Comment by Matumio 13 hours ago
When you say "cross-entropy loss" people without stats background go to Wikipedia, take a glance, and adjust their mental model to "inscrutable magic".
Thinking of the main difference as the trade-off in how much CPU, memory and storage is allowed is not really wrong.
The part that is wrong is to think of gzip as a method that might reach similar complexity or generalization. And more importantly, to ignore the advanced way how training data gets curated or generated for (instructed, chain-of-thought) LLMs. But even then. The mental model that the LLM's goal is text compression is not wrong. The question to ask next is what kind of text it is expecting to compress.
Comment by montebicyclelo 12 hours ago
I do think when making these comparisons, it is worth emphasising that neural nets are really different. E.g. I used to see people equating LLMs to n-gram models, etc. which is overly simplistic, (especially in the early days when the models weren't as good).
Comment by northlondoner 7 hours ago
Learning being a compression is also recently proposed as Gibbs compression proposition.
See Gibbs randomness-compression proposition https://arxiv.org/abs/2505.23869v5
Comment by teh64 3 hours ago
Video from tsoding where he implements the algorithm: https://www.youtube.com/watch?v=9n39SbRPXKQ
Comment by tromp 15 hours ago
Comment by gkbrk 15 hours ago
Comment by computably 14 hours ago
Comment by Ohentis 2 hours ago
Comment by londons_explore 14 hours ago
Comment by anax32 14 hours ago
Comment by alienbaby 5 hours ago
Comment by nomel 4 hours ago
The goal would be to find the minimum model that, with a fixed seed, would exactly reproduce your text.
Comment by anothereng 2 hours ago
Comment by Ohentis 2 hours ago
Comment by aghilmort 8 hours ago
Comment by colinmarc 12 hours ago
Comment by berkes 14 hours ago
Some models are reproducible, in that the same prompt will generate the same output. Say that we could wire up such a model to generate some code.
In that case, we could create a prompt that generates, say, an entire codebase, or a large piece of text. The prompt (or really, the tokens) would then be the compressed version of the codebase or the text.
I am not talking about an "AI agent", but really a model that we call in a reproducible manner. Preferably one call, with one prompt. An agent could just run `git clone` to "decompress" a codebase, which conflates the idea of compression. If that were compression, then the "compressed version of the git kernel" would be a single line of text: `git clone https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...`. I am really talking about having an LLM re-generate text based on a prompt.
Does that make sense? I can imagine that this is highly impractical and inefficient. But would this count as "compression" at all?
Comment by evgpbfhnr 14 hours ago
Comment by stackbutterflow 13 hours ago
Comment by eru 13 hours ago
A large language model itself (the network) give you the probabilities for the next token given some prefix of tokens so far. You can use arithmetic coding to go from these probabilities to a deterministic compression / decompression algorithm.
When you use an LLM to generate text, you sample from that probability distribution. You can use a true random sample. Or you can make it trivially deterministic by using a seeded pseudo-random-number-generator or you just pick the highest probability each time. But that's all a red herring; really, what you want is arithmetic coding.
Comment by meindnoch 11 hours ago
A chat?
>I am not talking about an "AI agent", but really a model that we call in a reproducible manner.
An LLM is just as deterministic as any other computer program. For identical inputs (which includes the PRNG seed) it produces identical outputs.
>compressed version of the git kernel
The git kernel, got it.
>But would this count as "compression" at all?
Yes. The decompressor is several tens of gigabytes though.
Comment by foldr 11 hours ago
This is not really true in practice because of multi-threading and out-of-order execution. Mathematically equivalent orderings of operations are not equivalent when dealing with floating point values, so most practical LLM implementations end up being non-deterministic.
Comment by MarkusQ 8 hours ago
Comment by foldr 8 hours ago
Comment by vasco 12 hours ago
Like when you click the Calculator button on your android, it wouldn't actually exist yet, your click actually prompts it into existence. But naively that has problems because you don't want a different UI every time. There's something to your idea.
Comment by StilesCrisis 11 hours ago
Comment by dist-epoch 12 hours ago
https://imalogic.com/blog/2024/06/03/image-compression-decom...
Comment by flyinglizard 13 hours ago
Comment by Tornhoof 15 hours ago
Comment by adityaathalye 13 hours ago
Viz. if Language is compression (of thought / culture / the tacit je ne sait quois of being-to-being communication etc.), then definitionally, Language Modelling must also be Compression.
Except, language is an arbitrarily lossy compressor, who's "compression-prediction equivalence" is indeterminate and unstable, because Language co-evolves constantly; both as a function of or response to culture, as well as an influencer of culture.
So, the subjective-objective goodness of Language Models (of any kind of language) would be, at best, upper-bounded by the compression-prediction equivalence of the Languages corpus itself. And that is assuming the language corpus is perfect in every way---it captures all knowledge expressible by language and it is always in-sync with live evolution of all language expression and evolution (i.e. LLM training is not a batch job, but a real-time present continuous process).
For example, to my layperson eyes, the mathematical language of proofs actively weeds out ambiguity of subjective interpretation. Ideally, a proof ought to lead to the exact same conclusion on every single reading by any reader who can follow the steps. A proof also holds only if the rest of the formal, explicit, inviolable, internally-consistent set of axioms and results holds.
So it stands to reason that mathematical prose of proofs, being optimised as mechanical procedure of taking an open question to a deterministically closed solution, has better odds of approximating the tacit aspects of mathematical derivation.
Which makes an LLM able to construct a mathematical proof, which is mind-melting to say the least.
However, I wonder, can LLMs dream of mathematical sheep?
Comment by cestith 9 hours ago
It makes a lot of sense why dictionary-based compression is named the way it is. A shorter symbol is used to store information that would take more symbols in the uncompressed corpus, if the shorter symbol hadn't been assigned to represent it. That's in a way just what an actual dictionary on your English professor's shelf does. The big difference is your compressor is coining new short symbols all the time.
Comment by IAmBroom 2 hours ago
Languages add new words constantly, albeit slower than a computer does compressing a new file.
Comment by kazinator 7 hours ago
Comment by js98 6 hours ago
Comment by _def 8 hours ago
Comment by modin 13 hours ago
Comment by mentalgear 14 hours ago
Comment by networked 14 hours ago
Comment by Sesse__ 14 hours ago
Comment by networked 13 hours ago
gzipt \
--corpus data/tinyshakespeare.txt \
--prompt $'MENENIUS:\n' \
--length 200 \
;
MENENIUS:
MtLUMSeptuttyyyxyxyxyxyvyyyxyxyxyxyvyyyxyxyxyxywyvzyxyxyx
yyxyyyxyxyxyxyxPlyxyxyxyxyxyxyxyxyxtoxzfTUS.zxzzzyzzzvzzz
vzzzxvzyvyxyxyxyvyxyxyxyvy--,Vdvyxyxyxyxyxyxyxyxyxxy!zFlx
zzyyxyxyxyvyxyxyxyvyySPffuyuy
Line breaks added. This looks roughly optimized for the most repetitive Burrows-Wheeler transform (https://en.wikipedia.org/wiki/Burrows%E2%80%93Wheeler_transf...). Why are they runs of alternating symbols and not one symbol?Zstandard produces whitespace with the occasional letter thrown in. To quote MiMo: "As you can see, zstd does not speak Shakespeare. ... zstd encodes a run of one repeated byte as a near-free run-length sequence, and space and newline are the cheapest literals in the corpus: ten newlines cost about the same to append ten bytes of genuine corpus text and less than nonsense does."
Comment by maxidog 12 hours ago
Comment by networked 12 hours ago
This was the main change for bzip2:
@@ -33,19 +34,16 @@ def candidate_lengths(
level: int = 9,
pool: ThreadPoolExecutor | None = None,
) -> list[int]:
- """Compressed length of ``context + seq`` for each seq, sharing the context.
+ """Compressed length of ``context + seq`` for each seq.
- Compresses ``context`` once into a ``compressobj``, then clones its encoder
- state per candidate and feeds only that candidate. Identical to
- ``len(zlib.compress(context + seq, level))`` for each seq, but the expensive
- match search over ``context`` happens a single time.
+ Unlike ``zlib``'s ``compressobj``, Python's ``BZ2Compressor`` cannot be
+ snapshotted mid-stream, and bzip2's move-to-front + Huffman stages see the
+ whole block, so every candidate recompresses the full context. Threads
+ still scale because ``bz2`` releases the GIL.
"""
- base = zlib.compressobj(level)
- head = len(base.compress(context))
def length_for(seq: bytes) -> int:
- clone = base.copy()
- return head + len(clone.compress(seq) + clone.flush(zlib.Z_FINISH))
+ return len(bz2.compress(context + seq, level))
if pool is not None:
return list(pool.map(length_for, sequences))Comment by jrmg 9 hours ago
(I honestly don’t know is gzip does something different when presented with two chunks as opposed to one, or, if it does, if bz2 has equivalent behaviour - but the difference in the code did stand out to me, and it does seem related to ‘extending the token sequence’)
Comment by networked 7 hours ago
We can test it by going back to zlib:
def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(zlib.compress(context + seq, level))
At temperature zero, this outputs the same sample as commit 3734bf6, the most recent commit upstream: MENENIUS:
'Though all at once cannq
MARCIUS:
I'll fight
'Though all at once cannq
MARCIUannq
MARCIUS:
I'll fight
'Though
AUFIDIUS:
If I fly, Marci
AUFIDIUS:
If I fly, Marci
AUFID
AUFIDIUS:
If
If I fly
I also tried LZMA for good measure: def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(lzma.compress(context + seq))
The sample at temperature zero: MENENIUS:
'Th
A carbuncle enti
, as big as thou
A aa
This is followed by a lot of whitespace.python-lz4 gives you all newlines after the prompt. I tried debugging it, and the compressed length of different candidate seqs is the same.
Comment by jeremyjh 12 hours ago
Comment by elendilm 13 hours ago
Compression is a property of language.
A seemingly simple sentence like "I had lunch" has enormous amount of information compressed inside it.
The word lunch is a compressed form of "having food at noon" while "noon" in turn is a compressed form of "Sun's position against Earth's rotation" and so on and so forth.
Every sentence has layers of compressed sentences. How many layers one chooses to decompress is up to the person.
Comment by jcattle 7 hours ago
Comment by ronfriedhaber 9 hours ago
Comment by nelox 14 hours ago
Comment by northlondoner 7 hours ago
Comment by dominotw 9 hours ago
Comment by DonHopkins 11 hours ago
In the 2023 discussion of "Demoscene accepted as UNESCO cultural heritage in The Netherlands" I posted a transcript from a video of Will Wright discussing the demo scene:
https://news.ycombinator.com/item?id=36599415
Will Wright Discusses the Demoscene:
https://www.youtube.com/watch?v=m7iuFVmTJus
>You can take any piece of content in the game, and imagine an algorithmic solution to it. Or also, you know, a way that the player could customize that object of thing.
>There's this group in Europe called the Demoscene that make these very elaborate demos for a computer that fit into very tiny little memory blocks, you know like 64K of memory, and you run the thing, and in fact it algorithmically generates about 100 megabytes worth of data, you know these rich 3D environment, generated music, generated wave files, generated animation.
>And they're developing techniques to generate, you know, huge amounts of interesting data, with very very simple, elegant, compression algorithms.
>And this is a skill that game developers used to have, back in the 8-bit days. That was the only ways to do a game like Karateka(?), was to find all these little tips and tricks to compress things and generate them algorithmically.
>But since the CD-ROM came out, and very cheap hard drives, storage is cheap, so basically we've lost that skill set, and now we attack all those problems with brute force. I think we've lost something by dropping that skill set.
[...]
https://news.ycombinator.com/item?id=36613058
[...] Here's a simple low-tech pre-LLM example that shows the equivalence of compression and procedural content generation:
Take a huge text file of HN postings, and compress it with gzip or compress or some other robust compression algorithm. The better the algorithm, the more the output will look like random noise. Then slice the compressed file in half, and replace the second half with random numbers. Then uncompress it. You'll find that at the point you sliced it, it keeps on writing out almost plausible text for a while, consisting of highly probably snippets of commonly encountered words and phrases, then goes downhill towards incoherence. It's not as coherent or confident as an LLM, but the point is to show how low the bar is for using compression for procedural content generation.
LLMs are essentially a form of compression of the world's knowledge or whatever they're trained on, not just word frequencies or pixel patterns, but also concepts and ideas. [...]
Comment by mohd_rafay 9 hours ago
Comment by corbinvachal 7 hours ago
Comment by kindkang2024 10 hours ago
Comment by lotus_uk 7 hours ago
Comment by greengemz 9 hours ago
Comment by fr2029 14 hours ago
Comment by 0x20cowboy 14 hours ago
Comment by relevant_stats 14 hours ago
Some will say that I should 'judge the idea, not the form'.
But if the author didn't find enough strength to write alone a short ~700 words summary about his work, it means he himself isn't that interested or enthusiastic about it. Why should others bother then? Particularly since low-effort like that signals possibility the whole work is superficial and derivative.
Comment by marand23 13 hours ago
Comment by relevant_stats 9 hours ago
Comment by DonHopkins 11 hours ago
Your claim that suspected AI assistance proves the author isn't interested -- and therefore that the work is probably superficial -- is unsupported.
The article presents a working experiment, explains why naive decoding fails, describes the beam-search fix, and links the code.
Dismissing all that with presumptuous personal speculation and banal boilerplate drive-by anti-AI snark adds absolutely nothing to the conversation -- and that is intrinsically poor form.
You couldn't even find enough strength to criticize anything beyond the form, while your own form is lackluster.
Ironically, an LLM could have written your comment and improved its form without losing anything distinctive.
Comment by relevant_stats 9 hours ago
and simultaneously you write that 'my form is lackluster' and that 'an LLM could have written your comment and improved its form without losing anything distinctive'. We are having ourselves a small contradiction, aren't we.
Be my guest, enjoy chatbot writing and drowning in slop. But don't encroach upon my freedom to protest it.
Comment by DonHopkins 9 hours ago
Your claim that suspected AI assistance proves the author isn't interested -- and therefore that the work is probably superficial -- is unsupported.
So I'm encroaching but you're only protesting, huh? I also have the freedom to ironically protest the poor form of your inability to criticize ideas, as well as your poorly formulated unsupportable ideas.
Comment by relevant_stats 9 hours ago
Pointing out someone's contradiction is now being 'critical of form'? And someone's contradicting themselves is 'ironic'?
Now I'm not even sure you know the meaning of words you use. EOT from me.
Comment by elendilm 7 hours ago
If you then use AI to write a product page documentation and proof read it for correctness, would that constitute to signaling that the whole effort is superficial?
Some work may be left to AI while you focus on the more important aspects of the work.
Surprisingly people have got it completely backwards where they want AI to generate code and humans to write documentation.
Comment by bob1029 15 hours ago
The fact that gzip is relatively fast should be your first clue that something important is missing.
Gzip is great at predicting the next token for one very specific narrative. LLMs can predict next tokens for entire universes of narratives. Searching for the correct next token across this space scales ~quadratically with the input size. Gzip scales linearly. I can gzip a one terabyte file. Imagine feeding that much into an LLM. These are wildly different animals that happen to overlap in a very small way. Equating compression to intelligence looks increasingly silly to me.
If we must compare language models to compression, they are much more like jpeg and mp3 than they are gzip and flac. I can go fuck with a jpeg file pretty severely at the bitstream level and still have something resembling performance on the other side. Gzip cannot remotely approach this.
Comment by Retr0id 15 hours ago
In part because gzip only has a 32KiB window size, and I think it'd be at least quadratic within that window if you were going for optimal compression.
Comment by bob1029 14 hours ago
Show me an LLM that can run at 300 megabytes per second. Even dedicated ASICs with weights burned in will never move this fast.
Comment by pishpash 13 hours ago
Comment by Sesse__ 15 hours ago
Comment by fedeb95 14 hours ago
Comment by amelius 15 hours ago
Comment by magicalhippo 15 hours ago
Quite well. This project[1], by Fabrice Bellard of ffmpeg fame, is quite old in AI years and uses an ancient LLM, but still beats xz by a solid margin.
Comment by amelius 13 hours ago
Comment by magicalhippo 13 hours ago
A challenge as I understand it is reproducibility.
Normal LLM runtimes aren't typically fully reproducible even with same random seeds for distribution sampling, due to floating-point numbers, batching and such.
Though averaging over many runs could alleviate that I suppose.
While it would measure some aspects of intelligence, I'd argue it fails to capture other, more creative aspects.
Comment by segmondy 11 hours ago