<?xml version="1.0" encoding="UTF-8" ?>
<rss version="2.0">
    <channel>
      <title>📖 llm-tracker</title>
      <link>https://llm-tracker.info</link>
      <description>Last 10 notes on 📖 llm-tracker</description>
      <generator>Quartz -- quartz.jzhao.xyz</generator>
      <item>
    <title>MI300X MoE Training</title>
    <link>https://llm-tracker.info/MI300X-MoE-Training</link>
    <guid>https://llm-tracker.info/MI300X-MoE-Training</guid>
    <description>075 - TRL + Custom megablocks-hip fork gets to step 0 at least Sample Blog - GPT2 training: https://rocm.blogs.amd.com/artificial-intelligence/megablocks/README.</description>
    <pubDate>Tue, 17 Mar 2026 17:12:13 GMT</pubDate>
  </item><item>
    <title>Untitled 5</title>
    <link>https://llm-tracker.info/Untitled-5</link>
    <guid>https://llm-tracker.info/Untitled-5</guid>
    <description></description>
    <pubDate>Tue, 17 Mar 2026 17:12:13 GMT</pubDate>
  </item><item>
    <title>RTX PRO 6000</title>
    <link>https://llm-tracker.info/RTX-PRO-6000</link>
    <guid>https://llm-tracker.info/RTX-PRO-6000</guid>
    <description>llama.cpp llama2-7b 600W ❯ build/bin/llama-bench -m /models/llm/gguf/llama-2-7b.Q4_0.gguf -fa 1 ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no ggml_cuda_init: found 1 CUDA devices: Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.</description>
    <pubDate>Tue, 09 Sep 2025 16:42:35 GMT</pubDate>
  </item><item>
    <title>ms-swift</title>
    <link>https://llm-tracker.info/ms-swift</link>
    <guid>https://llm-tracker.info/ms-swift</guid>
    <description>https://www.notion.so/Ascend_Doc-2180dc3bf51680989af4cd4eee46acdd Training (ms-swift) https://swift.readthedocs.io/en/latest/BestPractices/NPU-support.</description>
    <pubDate>Fri, 08 Aug 2025 17:19:51 GMT</pubDate>
  </item><item>
    <title>Strix Halo</title>
    <link>https://llm-tracker.info/_TOORG/Strix-Halo</link>
    <guid>https://llm-tracker.info/_TOORG/Strix-Halo</guid>
    <description>Those looking for my testing code: https://github.com/lhl/strix-halo-testing I will probably consolidate and redirect to the Github repo at some point.</description>
    <pubDate>Fri, 08 Aug 2025 15:07:10 GMT</pubDate>
  </item><item>
    <title>AI Server</title>
    <link>https://llm-tracker.info/AI-Server</link>
    <guid>https://llm-tracker.info/AI-Server</guid>
    <description>For best price/perf, Dual Socket EPYC ROME is probably the way to go. If you have the cash: Cheap dual core 9004 chips. The going rate for a 9334 QS/ES chip is 600atm,soyoucouldgetapairfor1200 and should give you about 400GB/s For dual socket you’re probably looking at a Gigabyte MZ73-LM1 or AsRock Rack TURIN2D16-2T - looks like those aren’t going for about $1000-1300 It’s about $1500-1800 for 384GB (24x16GB or 12x32GB DDR5-4800 ECC); 30% more for DDR5-5600 but maybe worth it if you’re going to drop in a 9005 upgrade at some point 250fora6U&quot;mining&quot;GPUservercase(AliExpress),shouldbeabletoeasilyfit4GPUsonrisers(notanissueforinferencing)−say200 for good quality risers $400 for a good 1300-1500W power supply Either way you’ll want GPUs, depending on your price: CostGPUMemoryMBWFP16 TFLOPSNotes$8200Nvidia RTX PRO 600096GB1.</description>
    <pubDate>Mon, 16 Jun 2025 18:14:44 GMT</pubDate>
  </item><item>
    <title>Best Courses</title>
    <link>https://llm-tracker.info/Best-Courses</link>
    <guid>https://llm-tracker.info/Best-Courses</guid>
    <description>Practical mlabonne’s LLM Course https://github.com/mlabonne/llm-course Mastering LLMs https://hamel.dev/blog/posts/course/ Evals https://www.youtube.com/playlist?list=PLgIaq8VgndJvt-HKMHPXehyJNNXQsAVHD https://hamel.</description>
    <pubDate>Fri, 06 Jun 2025 06:32:52 GMT</pubDate>
  </item><item>
    <title>Quant JA MT-Bench Comparison</title>
    <link>https://llm-tracker.info/Quant-JA-MT-Bench-Comparison</link>
    <guid>https://llm-tracker.info/Quant-JA-MT-Bench-Comparison</guid>
    <description>FP16 --- Scores for Model: shisa-ai/shisa-v2-llama3.1-405b --- Category gpt-4-0613 gpt-4-turbo gpt-4.1-2025-04-14 gpt-4.1-mini-2025-04-14 gpt-4o ----------- ------------ ------------- -------------------- ------------------------- -------- coding 9.</description>
    <pubDate>Fri, 06 Jun 2025 01:48:09 GMT</pubDate>
  </item><item>
    <title>AMD GPUs</title>
    <link>https://llm-tracker.info/howto/AMD-GPUs</link>
    <guid>https://llm-tracker.info/howto/AMD-GPUs</guid>
    <description>As of August 2023, AMD’s ROCm GPU compute software stack is available for Linux or Windows. It’s best to check the latest docs for information: https://rocm.</description>
    <pubDate>Fri, 23 May 2025 06:51:31 GMT</pubDate>
  </item><item>
    <title>LLM Inference Benchmarking Cheat‑Sheet for Hardware Reviewers</title>
    <link>https://llm-tracker.info/howto/LLM-Inference-Benchmarking-Cheat%E2%80%91Sheet-for-Hardware-Reviewers</link>
    <guid>https://llm-tracker.info/howto/LLM-Inference-Benchmarking-Cheat%E2%80%91Sheet-for-Hardware-Reviewers</guid>
    <description>NOTE: This document tries to avoid using the term “performance” since in ML research the term performance typically refers to measuring model quality/capabilities.</description>
    <pubDate>Sun, 18 May 2025 05:39:09 GMT</pubDate>
  </item>
    </channel>
  </rss>