CodeInvasion

@CodeInvasion@sh.itjust.works

This is a remote user, information on this page may be incomplete. View at Source ↗

⁨Comment⁩ on ⁨I'm gonna die on this hill or die trying⁩ ⁨⁨3⁩ ⁨weeks⁩ ago⁩:
The tokenizer is capable of decoding spaceless tokens into compound words following a set of rules referred to as a grammar in Natural Language Processing (NLP). I do LLM research and have spent an uncomfortable amount of time staring at the encoded outputs of most tokenizers when debugging. Normally spaces are not included.

There is of course a token for spaces in special circumstances, but I don’t know exactly how each tokenizer implements those spaces. So it does make sense that some models would be capable of the behavior you find in your tests, but that appears to be an emergent behavior, which is very interesting to see it work successfully.

I intended for my original comment to convey the idea that it’s not surprising that LLMs might fail at following the instructions to include spaces since it normally doesn’t see spaces except in special circumstances. Similar to how it’s unsurprising that LLMs are bad at numerical operations because of how the use Markov Chain probability to each next token, one at a time.
⁨Comment⁩ on ⁨I'm gonna die on this hill or die trying⁩ ⁨⁨3⁩ ⁨weeks⁩ ago⁩:
This is because spaces typically are encoded by model tokenizers.

In many cases it would be redundant to show spaces, so tokenizers collapse them down to no spaces at all. Instead the model reads tokens as if the spaces never existed.

For example it might output: thequickbrownfoxjumpsoverthelazydog

Except it would actually be a list of numbers like: [1, 256, 6273, 7836, 1922, 2244, 3245, 256, 6734, 1176, 2]

Then the tokenizer decodes this and adds the spaces because they are assumed to be there. The tokenizer has no knowledge of your request, and the model output typically does not include spaces, hencr your output sentence will not have double spaces.
⁨Comment⁩ on ⁨Are Cars Just Becoming Giant Smartphones on Wheels?⁩ ⁨⁨1⁩ ⁨month⁩ ago⁩:
Thanks, it’s definitely a good escape from the code, that’s for damn sure.
⁨Comment⁩ on ⁨Are Cars Just Becoming Giant Smartphones on Wheels?⁩ ⁨⁨1⁩ ⁨month⁩ ago⁩:
Old cars are work for sure, but if you are willing to learn it’s not bad.

I have a 2007 Mustang. I’ve replaced the entire front suspension, rear differential, and paid an upholsterer to replace the convertible top. I upgraded the radio and put in a 10inch touch screen with Wireless carplay and integrated backup camera. Next up is dropping the trans to replace the clutch plates, throw-out bearing, resurface the flywheel, and replace the rear main seal on the engine while I’m down there because the flywheel is rusty accumulates a thin layer every morning that makes a grinding noise for 30 seconds until it grinds off.

It definitely doesn’t just work like a new car, but since I do the work myself it also doesn’t cost me much.
⁨Comment⁩ on ⁨Elon Musk and Taylor Swift can now hide details of their private jets/// Private aircraft owners can now ask the FAA to keep their registration information out of the public eye.⁩ ⁨⁨6⁩ ⁨months⁩ ago⁩:
Absolutely air traffic in the sky should be identified. There is no problem with that, but it’s the idea that it is too easy to find out everything about an aircraft owner by simply seeing the number on their tail.

The rich guys obfuscate that info with shell corps to own the aircraft.

Shouldn’t everyone have the right to the same level of privacy regardless of how much money they have?
⁨Comment⁩ on ⁨Elon Musk and Taylor Swift can now hide details of their private jets/// Private aircraft owners can now ask the FAA to keep their registration information out of the public eye.⁩ ⁨⁨6⁩ ⁨months⁩ ago⁩:
No you cannot. You cannot easily find someone’s address from looking at their plate. You need more information, or to do some advanced searching. It is simply not the same.
⁨Comment⁩ on ⁨Elon Musk and Taylor Swift can now hide details of their private jets/// Private aircraft owners can now ask the FAA to keep their registration information out of the public eye.⁩ ⁨⁨6⁩ ⁨months⁩ ago⁩:
It is different because you typically need to know the municipality I live in first.

Also the registration allows anyone to track me anytime I fly.

How would you feel if you had a public gps transponder on your car publicly showing who you, where you are, and where you live? Also what if you are required to plaster that registration number on the side of your vehicle in large letters that can be seen from a block away?

It’s a massive invasion of personal privacy.
⁨Comment⁩ on ⁨Elon Musk and Taylor Swift can now hide details of their private jets/// Private aircraft owners can now ask the FAA to keep their registration information out of the public eye.⁩ ⁨⁨6⁩ ⁨months⁩ ago⁩:
Shitposters ride for free
⁨Comment⁩ on ⁨Elon Musk and Taylor Swift can now hide details of their private jets/// Private aircraft owners can now ask the FAA to keep their registration information out of the public eye.⁩ ⁨⁨6⁩ ⁨months⁩ ago⁩:
This is actually most helpful to the little guys that own $20,000 airplanes.

I have a small airplane and it’s always bothered me that my name and address are publicly accessible through the FAA registry.

Most pilots I know are careful about photos they publish online showing their tail number printed in large bold letters on either side of the aircraft. This registration number can be entered into websites like flightaware.com and someone is literally two clicks from seeing my full name and home address.
⁨Comment⁩ on ⁨Brian Eno: “The biggest problem about AI is not intrinsic to AI. It’s to do with the fact that it’s owned by the same few people”⁩ ⁨⁨7⁩ ⁨months⁩ ago⁩:
Well, OpenAI hast clearly scraped everything that is scrap-able on the internet. Copyrights be damned. I haven’t actually used Deep seek very much to make a strong analysis, but I suspect Sam is just mad they got beat at their own game.

The real innovation that isn’t commonly talked about is the invention of Multihead Latent Attention (MLA), which is what drive the dramatic performance increases in both memory (59x) and computation (6x) efficiency. It’s an absolute game changer and I’m surprised OpenAI has released their own MLA model yet.

While on the subject of stealing data, I have been of the strong opinion that there is no such thing as copyright when it comes to training data. Humans learn by example and all works are derivative of those that came before, at least to some degree. This, if humans can’t be accused of using copyrighted text to learn how to write, then AI shouldn’t either. Just my hot take that I know is controversial outside of academic circles.
⁨Comment⁩ on ⁨Brian Eno: “The biggest problem about AI is not intrinsic to AI. It’s to do with the fact that it’s owned by the same few people”⁩ ⁨⁨7⁩ ⁨months⁩ ago⁩:
Yah, I’m an AI researcher and with the weights released for deep seek anybody can run an enterprise level AI assistant. To run the full model natively, it does require $100k in GPUs, but if one had that hardware it could easily be fine-tuned with something like LoRA for almost any application. Then that model can be distilled and quantized to run on gaming GPUs.

It’s really not that big of a barrier. Yes, $100k in hardware is, but from a non-profit entity perspective that is peanuts.

Also adding a vision encoder for images to deep seek would not be theoretically that difficult for the same reason. In fact, I’m working on research right now that finds GPT4o and o1 have similar vision capabilities, implying it’s the same first layer vision encoder and then textual chain of thought tokens are read by subsequent layers. (This is a very recent insight as of last week by my team, so if anyone can disprove that, I would be very interested to know!)