That’s an interesting suggestion. Given the variability in splitting possibilities, we will never be able to satisfy everyone. So, the ability to have pre-split text imported as is, bypassing our splitter, could be something to enable. On the other hand, it creates another control which has to be explained. The reality is that not much Japanese content is available pre-split so it would likely be a function which gets little usage.
Perhaps enabling users to customize the way they want texts split would be more useful. Some splitting instructions which would be passed to the LLM before it performs this task when generating the lesson from an import. We will think about this some more.
It is definitely personal preference how things are split. I find it myself when looking at content which has been split in different ways. At the moment our splitter is trying to identify words since we believe that is more important than focusing on the individual components. The fact is that you are recognizing the grouping or lack thereof already which is probably a good indication of your growing ability in the language. If the desire for split control is more about the type of words you want to include in your studies, maybe you have to let things go a bit. I have the same urges but in the end what matters is that you are improving and learning and growing your Japanese. Saving the same words in different forms and with different particles attached is actually fine. The big thing is to save them and keep moving. Those words that matter will reappear.
Splitting will never be perfect and the words you save and learn from a stats perspective will never be perfect either. That imperfection is just something we need to accept in language learning. That is the nature of it.
Having said that, whatever we can do to improve word splitting, transliteration, tts in Japanese, we will try and do. We can definitely be better. All I’m saying is that, at the same time, it works very well now from a learning Japanese perspective. I have been using it for Japanese for years. Of course it could be better but that is no reason to get turned off.
Thanks for the suggestion. I have passed it along.
考虑到人们喜好的不同,如果能让用户自己决定分割词汇的策略,那么就再好不过了。
如果是我自己的话,短语最好提前分割出来,形容词不需要加词尾,后缀。片假名的词汇也不需要分割,因为分割以后会一团糟,可能英语母语者对于片假名没什么感觉。
If Japanese was my main target language I’d be really upset. The sentence version audio does not match what is written and it is randomly generated guesses of the correct pronunciation. Sometimes even in Italian it adds weird sounds and words not in the original. AI has gotten worse not better in all languages even yelling the words but Japanese has the added negative of splitting.
I think the idea of having a toggle for automatic Ai splitting in settings would be great because it really does the job very well.
When using the ai version it more accurately reflects the unknown word count and when its not used it can differ as much as like 20% or 30% at times. This makes it hard to gauge lessons difficulty sometimes. I do agree its not the end of the world sure but its bothersome to have random blue lingq’s of things that just don’t make sense to even be blue in the first place like the example picture the op used.
全てを、俺が for example is just a waste of time to hit ignore when people just want to look at real unknown words.
Again, i know its never going to be perfect for sure but Ai fixes most of this.
Are those in recently imported lessons? I thought we had removed most of the splits that attached particles to words. That shouldn’t still be happening. All imports are split with AI.
@Vanguard96 Can you provide some examples and links so we can see what’s going on with the TTS? The AI voices do sometimes act strangely but with some examples we can try to correct these issues.