What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks Paper • 2504.07825 • Published Apr 10, 2025 • 1
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges Paper • 2605.23970 • Published May 13 • 2
Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential Paper • 2610.05541 • Published 8 days ago • 2
view article Article Making LLMs Smaller Without Breaking Them: A GLU-Aware Pruning Approach oopere • Nov 24, 2024 • 23
view article Article Frontier-Assisted Single-Prompt Disposable Risk Assessment darkc0de • Aug 30 • 5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 217
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect Paper • 2403.03853 • Published Mar 6, 2024 • 66
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels Paper • 2406.17415 • Published Jun 25, 2024 • 1
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models Paper • 2608.17202 • Published Aug 17 • 1
Python Pre-compiled Binaries Collection These are python wheels that are built and saved in their completed format for specific cuda/torch combinations. Use in images or GPaaS instances. • 3 items • Updated 27 days ago • 1
Data Attribution of Emergent Misalignment with Persona Features Paper • 2608.11025 • Published Aug 11 • 1