AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP

Re-conversion notice (2026-09-23)

This artifact was re-converted on the AXQuant factory host with AXQuant 1.9.0 (commit 723eaeb1fb5b), which preserves the exact source tokenizer assets inside the converter and excludes the grafted MTP shard from the MLX-LM text-path view during requantization. Conversion integrity checks (byte-exact MTP/vision sidecars, tokenizer byte equality, greedy text smoke) were re-executed for this revision.

The previous revision (5ab39b24bfd7f65203be9b7823b1840486f58b6d) remains available in this repository's history, but users should re-download the latest revision of this model to get this build. The disclosures below still apply: this is a development repack, not a certified artifact.

Metadata correction (same day, follow-up revision): the packaged axquant_plan.json and axquant_manifest.json target_class was corrected from the allocator-inherited 8bit label to the actual product class MXFP4. Weight payloads are byte-identical to the first re-conversion revision; only those two metadata files (and this notice) changed.

Development repack. Not certified. No measured quality-parity or MTP-speed claim.

This is an AXQuant MXFP4 requantization of the exact peculiar-ragdoll/Tiel-Coder-35B-A3B-MLX-oQ6e-MTP checkpoint, revision 88625754ac91b542280a5602239ce6b2166366f0. The source is already mixed-precision oQ6e, not original BF16. Public MLX APIs dequantize its weights and repack the language trunk using an AXQuant manual recipe. This adds quantization error and does not restore the original full-precision weights. Upstream benchmark results and coding-performance claims do not apply to this repack. No upstream importance matrix, calibration corpus, or quantizer implementation is imported.

Artifact

Property Value
AXQuant version 1.9.0
AXQuant commit 723eaeb1fb5b
Architecture Qwen3.5-class 35B-A3B MoE; qwen35-moe-v1
Language trunk Native MLX MXFP4, group size 32
Protection floors Embeddings/routers at least 8-bit; norms/LM head BF16
Measured main BPW 4.634223
Measured total BPW 4.901270
Weight bytes 22026200211
MTP 785 source tensors preserved in mtp.safetensors
Vision Source vision tensors preserved in vision.safetensors
Previous revision 5ab39b24bfd7f65203be9b7823b1840486f58b6d
Certification None; architecture-prior manual allocation

The protected MTP and vision tensor payloads were compared byte-for-byte with the pinned source. The source tokenizer and Sharp chat template are preserved. Basic deterministic text-generation smoke tests ran with MLX-LM on the factory host. These checks establish conversion integrity and basic loadability, not model quality, vision capability, MTP acceptance, or acceleration. AX Engine execution was not measured for this artifact in this campaign.

Download and basic text generation

python -m pip install mlx-lm huggingface_hub
hf download AutomatosX/AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP
mlx_lm.generate --model ./AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP --max-tokens 256 --prompt "Write a Python add function."

Use an Apple Silicon Mac with sufficient unified memory. Pin the published Hub commit for reproducible deployments. The tested packaging environment used MLX-LM 0.31.3. MXFP4 here is the native MLX Apple format, not NVIDIA NVFP4.

MTP and multimodal scope

The -MTP suffix means the head is packaged, not that acceleration is certified or enabled by the text command above. Stock MLX-LM text generation does not use the separate MTP sidecar. A compatible sidecar-aware runtime is required, and runtime-specific MTP enablement must be validated separately. No MTP grafting or new training was performed. The upstream MTP provenance remains that of the exact source repository. Vision weights are retained but this campaign does not claim vision inference support.

Provenance and license

The upstream model card declares MIT. Follow its license terms and usage information. Credits: peculiar-ragdoll for the exact source build and template, Ornith for the model lineage, and the additional upstream contributors identified in the source model card. AXQuant uses public MLX conversion APIs. See axquant_plan.json, axquant_quantizer_execution.json, the protected-sidecar manifests, and axquant_manifest.json for allocation and file bindings.

Downloads last month
821
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP

Collections including AutomatosX/AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP