Instructions to use FahimehBahman/PolyTAO-Student with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FahimehBahman/PolyTAO-Student with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("FahimehBahman/PolyTAO-Student") model = AutoModelForSeq2SeqLM.from_pretrained("FahimehBahman/PolyTAO-Student", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- PolyTAO Student: Knowledge-Distilled Transformer for Property-Conditioned Polymer Generation
- Research Objective
- Model Lineage
- Teacher Model
- Published Student Architecture
- Property Conditioning
- Knowledge Distillation
- Training Configuration
- Dataset
- Evaluation
- Chemical Validity
- Property Evaluation
- Distribution Analysis
- Experimental Pipeline
- Repository Files
- Loading the Model
- Loading the Property Projection
- Property-Conditioned Inference
- Intended Use
- Limitations
- Research Context
- Tools and Libraries
- Author
- Research Objective
PolyTAO Student: Knowledge-Distilled Transformer for Property-Conditioned Polymer Generation
PolyTAO Student is a property-conditioned Transformer model developed as part of my M.Sc. thesis, βKnowledge-Distilled Transformers for Property-Conditioned Polymer Generation.β
The project investigates whether a compressed Transformer student can retain useful generative and molecular-property information through knowledge distillation from a larger property-conditioned teacher.
The teacher was initialized from the publicly available hkqiu/PolyTAO-BigSMILES_Version checkpoint and subsequently trained with molecular-property conditioning.
The trained teacher was then frozen and used for feature-level knowledge distillation into smaller student configurations.
The checkpoint released in this repository corresponds to the training run configured with:
capacity_percent = 20
The released checkpoint contains 2 encoder layers and 12 decoder layers.
Research model: This checkpoint is provided for research and reproducibility purposes. It is not intended for production use or for making chemical or experimental decisions.
Research Objective
The central research question behind PolyTAO is:
How does Transformer compression affect chemical validity, property fidelity, and the generative behavior of property-conditioned polymer models?
The research investigates whether knowledge transferred from a larger teacher model can help a smaller student preserve useful internal representations and molecular-property information.
Multiple student configurations were explored during the research to analyze the relationship between model capacity and generative performance.
Model Lineage
The overall model-development pipeline is:
hkqiu/PolyTAO-BigSMILES_Version
β
βΌ
Property-Conditioned Teacher
β
β
β Feature-Level
β Knowledge Distillation
βΌ
PolyTAO Student Models
β
βΌ
Polymer Generation
β
βΌ
RDKit Evaluation
hkqiu/PolyTAO-BigSMILES_Version therefore serves as the starting checkpoint for the teacher rather than being the final student model published here.
Teacher Model
The teacher is based on a T5-style encoder-decoder Transformer architecture.
It was initialized from:
hkqiu/PolyTAO-BigSMILES_Version
using Hugging Face Transformers.
The teacher architecture uses:
Encoder layers: 12
Decoder layers: 12
d_model: 768
Attention heads: 12
A custom property-conditioning mechanism was added to the Transformer.
The teacher receives:
- a tokenized polymer representation, and
- a 15-dimensional molecular-property vector.
The property vector is mapped into the Transformer hidden space using a learned linear projection.
Conceptually:
15 Molecular Properties
β
βΌ
Linear Projection
β
βΌ
L2 Normalization
β
βΌ
Scaling (0.05)
β
βΌ
Property Embedding
β
βΌ
Encoder Hidden States
β
βΌ
Property-Conditioned
Encoder Representation
β
βΌ
Decoder
After training, the Transformer model, tokenizer, and learned property projection are saved separately.
Published Student Architecture
The checkpoint released in this repository corresponds to the training run configured with:
capacity_percent = 20
Inspection of the released checkpoint gives the following architecture:
Encoder layers: 2
Decoder layers: 12
d_model: 768
Attention heads: 12
The teacher and published student can therefore be summarized as:
| Component | Teacher | Published Student |
|---|---|---|
| Encoder layers | 12 | 2 |
| Decoder layers | 12 | 12 |
| Hidden dimension | 768 | 768 |
| Attention heads | 12 | 12 |
The released student substantially reduces encoder depth, while the decoder retains the 12-layer configuration.
The 20% designation refers to the capacity_percent setting used for this training run. It should not be interpreted as meaning that the complete released model contains exactly 20% of the teacher's total parameters.
The architecture reported in this Model Card reflects the actual configuration stored in the released checkpoint.
Property Conditioning
Both teacher and student use molecular-property conditioning.
The models are conditioned on the following 15 molecular properties:
MolWtHeavyAtomCountNHOHCountNOCountNumAliphaticCarbocyclesNumAliphaticHeterocyclesNumAliphaticRingsNumAromaticCarbocyclesNumAromaticHeterocyclesNumAromaticRingsNumHAcceptorsNumHDonorsNumHeteroatomsNumRotatableBondsRingCount
Normalized versions of the properties are used during training when available.
The conditioning process can be summarized as:
15-D Property Vector
β
βΌ
Linear Projection
β
βΌ
L2 Normalize
β
βΌ
Scale by 0.05
β
βΌ
Property Embedding
β
βΌ
Encoder Hidden States
β
βΌ
Conditioned Encoder
Representation
The learned property embedding is added to every position of the encoder hidden representation.
This allows molecular-property information to influence the representation used by the decoder during generation.
Knowledge Distillation
The student is trained using a frozen property-conditioned teacher.
For each training sample, teacher and student receive the same tokenized polymer representation and the same molecular-property vector.
The teacher produces a conditioned encoder representation.
The student is optimized both for the sequence-generation task and for reproducing the teacher's internal encoder representation.
Sequence Generation Loss
The student uses the standard sequence-to-sequence cross-entropy loss:
CE Loss
This trains the student to generate the target polymer sequence.
Feature-Level Distillation Loss
The student encoder hidden representation is compared with the corresponding frozen teacher representation.
Mean Squared Error is used for this feature-level distillation objective:
KD Loss = MSE(
Teacher Hidden Representation,
Student Hidden Representation
)
A learned teacher-to-student projection is used in the distillation pipeline where required.
The implemented student-training objective is:
Total Loss = CE Loss + Ξ± Γ KD Loss
with:
Ξ± = 0.5
The teacher remains frozen during student training.
Training Configuration
Teacher Training
The teacher training configuration includes:
| Parameter | Value |
|---|---|
| Starting checkpoint | hkqiu/PolyTAO-BigSMILES_Version |
| Optimizer | AdamW |
| Learning rate | 1e-5 |
| Batch size | 8 |
| Epochs | 3 |
| Property scale | 0.05 |
| Validation ratio | 0.1 |
| Precision | 32-bit |
The teacher was trained using PyTorch Lightning.
The learned teacher property projection was saved separately as:
property_proj.pt
Published Student Training Run
The released checkpoint corresponds to the run with:
capacity_percent: 20
learning_rate: 3e-5
teacher_checkpoint: teacher_final
Other training settings used in the student pipeline include:
| Parameter | Value |
|---|---|
| Optimizer | AdamW |
| Learning rate | 3e-5 |
| Batch size | 16 |
| Epochs | 3 |
| KD alpha | 0.5 |
| Property scale | 0.05 |
| Validation ratio | 0.1 |
| Random seed | 42 |
The implementation uses:
- PyTorch
- PyTorch Lightning
- Hugging Face Transformers
Dataset
The research dataset was derived from the Open Macromolecular Genome (OMG) polymer dataset.
The experimental dataset contains approximately 99,000 polymer samples, with each sample containing a polymer representation together with molecular-property information.
The preprocessing pipeline includes:
Polymer Dataset
β
βΌ
Cleaning & Filtering
β
βΌ
Molecular Property Processing
β
βΌ
Property Normalization
β
βΌ
Polymer Tokenization
β
βΌ
Model Input
The molecular descriptors used for conditioning and evaluation are calculated using RDKit.
The dataset itself is not distributed in this model repository.
Evaluation
The generated polymer representations are evaluated using complementary structural and property-based measures.
Evaluation considers:
- chemical validity of generated structures, and
- property fidelity of valid generated structures.
Chemical Validity
Generated sequences are parsed and sanitized using RDKit.
A generated structure is considered valid when RDKit can successfully construct and sanitize the corresponding molecule.
For the final evaluation of the published student checkpoint:
| Metric | Result |
|---|---|
| Generated samples | 500 |
| Valid samples | 250 |
| Chemical validity | 50.0% |
Only molecules that pass the RDKit validity check are included in downstream molecular-property evaluation.
Property Evaluation
For valid generated structures, the molecular properties are recalculated using RDKit and compared with the target conditioning properties.
Four complementary metrics are used.
Mean Absolute Error (MAE)
MAE measures the average absolute difference between generated and target property values.
Lower values indicate smaller average deviations.
Root Mean Squared Error (RMSE)
RMSE measures property error while giving greater weight to larger deviations.
Lower values indicate smaller errors.
Pearson Correlation
Pearson correlation measures the linear relationship between generated and target property values.
Higher positive correlation indicates stronger agreement in the direction of property variation.
Maximum Mean Discrepancy (MMD)
MMD is used to compare generated and reference property distributions.
Lower values indicate closer distributional agreement under the selected kernel.
However, MMD should not be interpreted in isolation.
A compressed generative model can produce a low MMD while generating a narrower or less diverse set of structures whose aggregate properties happen to resemble the reference distribution.
For this reason, MMD is interpreted together with:
- chemical validity
- MAE
- RMSE
- Pearson correlation
- qualitative distribution analysis
Distribution Analysis
Selected molecular-property distributions from the final evaluation are shown below.
Molecular Weight
Heavy Atom Count
Ring Count
These plots compare the distributions of valid generated structures with the corresponding reference/test distributions.
They provide a qualitative complement to the numerical evaluation metrics and help reveal differences that may not be captured by a single aggregate metric.
Experimental Pipeline
Open Macromolecular Genome
Dataset
β
βΌ
Data Preprocessing
β
βΌ
βββββββββββββββββββββββββββ
β Polymer Representation β
β + β
β 15 Molecular Properties β
ββββββββββββββ¬βββββββββββββ
β
βΌ
hkqiu/PolyTAO-BigSMILES_Version
β
βΌ
Property-Conditioned Teacher
β
βΌ
Teacher Hidden States
β
β
β Feature-Level
β Knowledge Distillation
βΌ
Property-Conditioned Student
β
βΌ
Polymer Generation
β
βΌ
RDKit Validation
β
βΌ
Valid Generated Molecules
β
βΌ
Property Recalculation
β
βΌ
βββββββββββββββββββββββββββ
β Chemical Validity β
β MAE β
β RMSE β
β MMD β
β Pearson Correlation β
ββββββββββββββ¬βββββββββββββ
β
βΌ
CapacityβPerformance Analysis
Repository Files
The released model includes the standard Transformer checkpoint together with the additional property-conditioning weights required by PolyTAO.
Important files include:
config.json
generation_config.json
model.safetensors
property_proj_student.pt
tokenizer.json
tokenizer_config.json
special_tokens_map.json
added_tokens.json
MolWt_hist.png
HeavyAtomCount_hist.png
RingCount_hist.png
README.md
model.safetensors
Contains the Transformer weights for the released student checkpoint.
property_proj_student.pt
Contains the learned linear projection used to map the 15-dimensional molecular-property vector into the Transformer's hidden space.
This file is required to reproduce the custom property-conditioning mechanism.
Loading the Model
The Transformer and tokenizer can be loaded directly from Hugging Face:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
MODEL_ID = "FahimehBahman/PolyTAO-Student"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSeq2SeqLM.from_pretrained(
MODEL_ID
)
model.eval()
print("Encoder layers:", len(model.encoder.block))
print("Decoder layers:", len(model.decoder.block))
print("Hidden dimension:", model.config.d_model)
For the released checkpoint, this should report:
Encoder layers: 2
Decoder layers: 12
Hidden dimension: 768
Loading the Property Projection
The custom property projection is stored separately from the standard Hugging Face Transformer checkpoint.
It can be downloaded and loaded as follows:
import torch
import torch.nn as nn
from huggingface_hub import hf_hub_download
MODEL_ID = "FahimehBahman/PolyTAO-Student"
projection_path = hf_hub_download(
repo_id=MODEL_ID,
filename="property_proj_student.pt"
)
property_proj = nn.Linear(
15,
model.config.d_model
)
state_dict = torch.load(
projection_path,
map_location="cpu"
)
property_proj.load_state_dict(state_dict)
property_proj.eval()
Property-Conditioned Inference
PolyTAO uses a custom property-conditioning mechanism.
Therefore, loading the Transformer checkpoint alone reproduces the Transformer architecture, but does not by itself reproduce property-conditioned generation.
The inference procedure requires:
Polymer Input
β
βΌ
Tokenizer
β
βΌ
Transformer Encoder
β
βΌ
Encoder Hidden States
β²
β
15 Molecular Properties
β
βΌ
Property Projection
β
βΌ
L2 Normalization
β
βΌ
Scale by 0.05
β
βΌ
Property Embedding
β
βΌ
Conditioned Encoder Representation
β
βΌ
Transformer Decoder
β
βΌ
Generated Polymer
A simplified implementation of the conditioning step is:
import torch
from transformers.modeling_outputs import BaseModelOutput
# properties:
# Tensor of shape [batch_size, 15]
# containing normalized molecular properties.
with torch.no_grad():
encoder_output = model.encoder(
input_ids=input_ids,
attention_mask=attention_mask,
return_dict=True
)
props = properties.clamp(-10.0, 10.0)
prop_emb = property_proj(props)
prop_emb = prop_emb / prop_emb.norm(
dim=-1,
keepdim=True
).clamp(min=1e-6)
prop_emb = 0.05 * prop_emb.unsqueeze(1)
conditioned_hidden = (
encoder_output.last_hidden_state
+ prop_emb
)
conditioned_encoder = BaseModelOutput(
last_hidden_state=conditioned_hidden
)
The conditioned encoder representation can then be supplied to the model decoder for autoregressive generation.
The property values must be normalized consistently with the preprocessing used during training.
Intended Use
PolyTAO Student is intended primarily for research involving:
- knowledge distillation
- Transformer model compression
- property-conditioned molecular generation
- polymer generative modeling
- efficient generative architectures
- representation-level knowledge transfer
- capacityβperformance analysis
The checkpoint can also serve as a research artifact for studying compressed Transformer models in molecular generation.
Limitations
This model has several important limitations.
- It is a research prototype, not a production model.
- Chemical validity does not imply synthesizability.
- Chemical validity does not imply chemical stability.
- Generated structures have not been experimentally validated.
- RDKit validity is a computational validity check rather than experimental verification.
- Property conditioning is limited to the molecular descriptors used during training.
- Property values must be normalized consistently with the original training preprocessing.
- Generation quality can depend substantially on decoding parameters.
- Distributional similarity does not guarantee structural diversity.
- A low MMD value alone does not establish superior generative quality.
- Property metrics are calculated only for structures that pass the validity check.
- The released checkpoint uses a custom conditioning mechanism outside the standard Hugging Face T5 forward interface.
- The model should not be used directly for safety-critical, medical, industrial, or experimental chemical decisions without independent expert validation.
Research Context
PolyTAO was developed as part of the M.Sc. thesis:
Knowledge-Distilled Transformers for Property-Conditioned Polymer Generation
The research investigates the effect of Transformer compression on property-conditioned polymer generation.
The work combines:
- Transformer-based generative modeling
- molecular-property conditioning
- feature-level knowledge distillation
- model compression
- computational molecular descriptors
- quantitative evaluation
Multiple student configurations were explored to study how model capacity influences chemical validity, property fidelity, and generated distributions.
The checkpoint published in this repository corresponds to the run configured with:
capacity_percent = 20
and contains:
2 encoder layers
12 decoder layers
d_model = 768
Tools and Libraries
The project uses:
- Python
- PyTorch
- PyTorch Lightning
- Hugging Face Transformers
- RDKit
- NumPy
- Pandas
- Scikit-learn
- Matplotlib
Author
Fahimeh Bahman
M.Sc. Data Science
Software Engineer & Machine Learning Researcher
Research interests include:
- Generative AI
- Transformer Models
- Knowledge Distillation
- Model Compression
- Machine Learning Systems
- AI for Software Engineering
- Downloads last month
- 32
Model tree for FahimehBahman/PolyTAO-Student
Base model
hkqiu/PolyTAO-BigSMILES_Version

