2 comments

  • ramigb 3 hours ago

    That's awesome I came from the jev in python thread. is the model shareable? do you have any more writing on this? would love to read more about it! cool demo anyways.

    • faangguyindia 3 hours ago

      Here you can see a 26B model hacked using a trained KV cache bank, responding like Gemma. The latency I am seeing is <127 ms.