My office computer has a Ryzen 7 5700, RX 580x, and 32gb of ram. Running ollama with deepseekv2 or llama3 is much slower than chatgpt in the browser. Same with my newer, more powerful home computer.

What kind of hardware do you need to run with comparable responsiveness to chatgpt? How much does it cost? Presuming such hardware is commercial, where do you find it?

  • JoYo@lemmy.ml
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    3 months ago

    It’s all dependent on VRAM. If you can load the distilled models with your GPU without maxing out your VRAM it will run just as fast as any server farm.

    RX 580x

    It looks like your video card only has 8 GB of VRAM. That will be your bottleneck.

  • zelifcam@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    edit-2
    3 months ago

    Install LocalAI and ensure it’s using acceleration. It’s one of the best solutions we have at the moment.

    Are you sure you’re not running these small models off of CPU and no acceleration? Because I’m running these small models pretty quickly. Nearly instant responses using a NVIDIA titanXP from a gaming rig I built in 2017 ish.

  • JASN_DE@lemmy.world
    link
    fedilink
    arrow-up
    4
    arrow-down
    3
    ·
    3 months ago

    You’d need basically a small server rack filled with datacenter GPUs. Expect mid to high 5 digit numbers.

    But: running smaller models on a typical gaming GPU is quite doable.

  • IHave69XiBucks@lemmygrad.ml
    link
    fedilink
    arrow-up
    1
    ·
    3 months ago

    you did not specify what type of model your trying to run. like deepseekr1 has various models if your trying to run the massive ones its not gonna work. You need to use a smaller model. I have a RX 6600 and run the 14b parameter model it does well.

    • IHave69XiBucks@lemmygrad.ml
      link
      fedilink
      arrow-up
      1
      ·
      3 months ago

      to be clear btw your CPU basically doesnt matter as far as i know. Just the GPU should be getting used any old CPU works. You CAN run it on a CPU but its gonna be very slow. But yeah the RX 6600 was decently cheap i got it for like 150$ so its not super expensive to run one of these models.