I found this subject interesting when I stumbled on it the other day.
The same amount of information does not necessarily generate/consume the same number of tokens in every language.
English is represented relatively efficiently by many commonly used tokenizers, but other languages can be chopped up into many more ...