Minkyoung Song
  • about
  • publications
  • blog (current)
  • cv
  • uncertainty
  • in-context-learning
  • model-compression
  • paper-notes
  • research-notes
  • Inference speed isn't just a model-size problem

    PagedAttention/vLLM, speculative decoding, and early-exit networks all speed up inference without shrinking a single parameter — three papers, one point.

    June 11, 2026

    2026 model-compression inference-acceleration efficiency research-notes

  • Why joint compression beats doing it in sequence

    APQ (Wang et al., CVPR 2020) shows what changes when architecture, pruning, and quantization are searched for jointly instead of tuned one after another.

    April 30, 2026

    2026 model-compression joint-optimization efficiency research-notes

  • Quantization: fewer bits per weight, without losing the model

    LLM.int8(), SmoothQuant, and GPTQ — three concrete answers to the same problem, outlier activations that make naive low-bit quantization of LLMs fail.

    March 21, 2026

    2026 model-compression quantization efficiency research-notes

  • Pruning: removing what a network doesn't need

    From Optimal Brain Damage to the Lottery Ticket Hypothesis — what the pruning literature actually removes, and why unstructured sparsity doesn't automatically buy speed.

    February 04, 2026

    2026 model-compression pruning efficiency research-notes

  • Is it the demonstrations, or the model?

    Ling, Zhao, Zhang et al. (NAACL 2024) decompose ICL's predictive uncertainty into a demonstration-selection component and a model component.

    November 14, 2025

    2025 uncertainty in-context-learning paper-notes

  • <
  • 1
  • 2
  • >
© Copyright 2026 Minkyoung Song. Powered by Jekyll with al-folio theme. Hosted by GitHub Pages. Photos from Unsplash.