One of the things OpenAI did to improve performance was to train an AI to determine how a human would rate an output, and use that to train the LLM itself. (Kinda like a GAN, now I think about it).
https://forum.effectivealtruism.org/posts/5mADSy8tNwtsmT3KG/...
But this process has probably gone as far as it can go, at least with current architectures for the parts, as per Amdahl's law.
One of the things OpenAI did to improve performance was to train an AI to determine how a human would rate an output, and use that to train the LLM itself. (Kinda like a GAN, now I think about it).
https://forum.effectivealtruism.org/posts/5mADSy8tNwtsmT3KG/...
But this process has probably gone as far as it can go, at least with current architectures for the parts, as per Amdahl's law.