Model repeating information and "spitting out" random characters

#12
by brazilianslib - opened

First of all, congratulations on the launch. Gemma 2 9B is, at least in my tests, the best model for PT-BR. Much better than much larger models.
However, problems are constantly happening, such as:

  1. Repeat information;
  2. "Spit" text infinitely;
  3. Place tags like "</start_of" at the end of your answer.

I am eagerly awaiting a solution.

Once again, I thank the entire Google Gemma team.

Google org

Would recommend you to use eager attention implementation

The same error even with eager attention and bf16.

Sign up or log in to comment