PolyTAO Student: Knowledge-Distilled Transformer for Property-Conditioned Polymer Generation

PolyTAO Student is a property-conditioned Transformer model developed as part of my M.Sc. thesis, β€œKnowledge-Distilled Transformers for Property-Conditioned Polymer Generation.”

The project investigates whether a compressed Transformer student can retain useful generative and molecular-property information through knowledge distillation from a larger property-conditioned teacher.

The teacher was initialized from the publicly available hkqiu/PolyTAO-BigSMILES_Version checkpoint and subsequently trained with molecular-property conditioning.

The trained teacher was then frozen and used for feature-level knowledge distillation into smaller student configurations.

The checkpoint released in this repository corresponds to the training run configured with:

capacity_percent = 20

The released checkpoint contains 2 encoder layers and 12 decoder layers.

Research model: This checkpoint is provided for research and reproducibility purposes. It is not intended for production use or for making chemical or experimental decisions.


Research Objective

The central research question behind PolyTAO is:

How does Transformer compression affect chemical validity, property fidelity, and the generative behavior of property-conditioned polymer models?

The research investigates whether knowledge transferred from a larger teacher model can help a smaller student preserve useful internal representations and molecular-property information.

Multiple student configurations were explored during the research to analyze the relationship between model capacity and generative performance.


Model Lineage

The overall model-development pipeline is:

hkqiu/PolyTAO-BigSMILES_Version
              β”‚
              β–Ό
   Property-Conditioned Teacher
              β”‚
              β”‚
              β”‚ Feature-Level
              β”‚ Knowledge Distillation
              β–Ό
     PolyTAO Student Models
              β”‚
              β–Ό
       Polymer Generation
              β”‚
              β–Ό
        RDKit Evaluation

hkqiu/PolyTAO-BigSMILES_Version therefore serves as the starting checkpoint for the teacher rather than being the final student model published here.


Teacher Model

The teacher is based on a T5-style encoder-decoder Transformer architecture.

It was initialized from:

hkqiu/PolyTAO-BigSMILES_Version

using Hugging Face Transformers.

The teacher architecture uses:

Encoder layers: 12
Decoder layers: 12
d_model:        768
Attention heads: 12

A custom property-conditioning mechanism was added to the Transformer.

The teacher receives:

  1. a tokenized polymer representation, and
  2. a 15-dimensional molecular-property vector.

The property vector is mapped into the Transformer hidden space using a learned linear projection.

Conceptually:

15 Molecular Properties
          β”‚
          β–Ό
   Linear Projection
          β”‚
          β–Ό
     L2 Normalization
          β”‚
          β–Ό
      Scaling (0.05)
          β”‚
          β–Ό
   Property Embedding
          β”‚
          β–Ό
Encoder Hidden States
          β”‚
          β–Ό
Property-Conditioned
Encoder Representation
          β”‚
          β–Ό
       Decoder

After training, the Transformer model, tokenizer, and learned property projection are saved separately.


Published Student Architecture

The checkpoint released in this repository corresponds to the training run configured with:

capacity_percent = 20

Inspection of the released checkpoint gives the following architecture:

Encoder layers:  2
Decoder layers: 12
d_model:         768
Attention heads: 12

The teacher and published student can therefore be summarized as:

Component Teacher Published Student
Encoder layers 12 2
Decoder layers 12 12
Hidden dimension 768 768
Attention heads 12 12

The released student substantially reduces encoder depth, while the decoder retains the 12-layer configuration.

The 20% designation refers to the capacity_percent setting used for this training run. It should not be interpreted as meaning that the complete released model contains exactly 20% of the teacher's total parameters.

The architecture reported in this Model Card reflects the actual configuration stored in the released checkpoint.


Property Conditioning

Both teacher and student use molecular-property conditioning.

The models are conditioned on the following 15 molecular properties:

  1. MolWt
  2. HeavyAtomCount
  3. NHOHCount
  4. NOCount
  5. NumAliphaticCarbocycles
  6. NumAliphaticHeterocycles
  7. NumAliphaticRings
  8. NumAromaticCarbocycles
  9. NumAromaticHeterocycles
  10. NumAromaticRings
  11. NumHAcceptors
  12. NumHDonors
  13. NumHeteroatoms
  14. NumRotatableBonds
  15. RingCount

Normalized versions of the properties are used during training when available.

The conditioning process can be summarized as:

15-D Property Vector
         β”‚
         β–Ό
   Linear Projection
         β”‚
         β–Ό
    L2 Normalize
         β”‚
         β–Ό
   Scale by 0.05
         β”‚
         β–Ό
 Property Embedding
         β”‚
         β–Ό
Encoder Hidden States
         β”‚
         β–Ό
Conditioned Encoder
   Representation

The learned property embedding is added to every position of the encoder hidden representation.

This allows molecular-property information to influence the representation used by the decoder during generation.


Knowledge Distillation

The student is trained using a frozen property-conditioned teacher.

For each training sample, teacher and student receive the same tokenized polymer representation and the same molecular-property vector.

The teacher produces a conditioned encoder representation.

The student is optimized both for the sequence-generation task and for reproducing the teacher's internal encoder representation.

Sequence Generation Loss

The student uses the standard sequence-to-sequence cross-entropy loss:

CE Loss

This trains the student to generate the target polymer sequence.

Feature-Level Distillation Loss

The student encoder hidden representation is compared with the corresponding frozen teacher representation.

Mean Squared Error is used for this feature-level distillation objective:

KD Loss = MSE(
    Teacher Hidden Representation,
    Student Hidden Representation
)

A learned teacher-to-student projection is used in the distillation pipeline where required.

The implemented student-training objective is:

Total Loss = CE Loss + Ξ± Γ— KD Loss

with:

Ξ± = 0.5

The teacher remains frozen during student training.


Training Configuration

Teacher Training

The teacher training configuration includes:

Parameter Value
Starting checkpoint hkqiu/PolyTAO-BigSMILES_Version
Optimizer AdamW
Learning rate 1e-5
Batch size 8
Epochs 3
Property scale 0.05
Validation ratio 0.1
Precision 32-bit

The teacher was trained using PyTorch Lightning.

The learned teacher property projection was saved separately as:

property_proj.pt

Published Student Training Run

The released checkpoint corresponds to the run with:

capacity_percent: 20
learning_rate: 3e-5
teacher_checkpoint: teacher_final

Other training settings used in the student pipeline include:

Parameter Value
Optimizer AdamW
Learning rate 3e-5
Batch size 16
Epochs 3
KD alpha 0.5
Property scale 0.05
Validation ratio 0.1
Random seed 42

The implementation uses:

  • PyTorch
  • PyTorch Lightning
  • Hugging Face Transformers

Dataset

The research dataset was derived from the Open Macromolecular Genome (OMG) polymer dataset.

The experimental dataset contains approximately 99,000 polymer samples, with each sample containing a polymer representation together with molecular-property information.

The preprocessing pipeline includes:

Polymer Dataset
      β”‚
      β–Ό
Cleaning & Filtering
      β”‚
      β–Ό
Molecular Property Processing
      β”‚
      β–Ό
Property Normalization
      β”‚
      β–Ό
Polymer Tokenization
      β”‚
      β–Ό
Model Input

The molecular descriptors used for conditioning and evaluation are calculated using RDKit.

The dataset itself is not distributed in this model repository.


Evaluation

The generated polymer representations are evaluated using complementary structural and property-based measures.

Evaluation considers:

  1. chemical validity of generated structures, and
  2. property fidelity of valid generated structures.

Chemical Validity

Generated sequences are parsed and sanitized using RDKit.

A generated structure is considered valid when RDKit can successfully construct and sanitize the corresponding molecule.

For the final evaluation of the published student checkpoint:

Metric Result
Generated samples 500
Valid samples 250
Chemical validity 50.0%

Only molecules that pass the RDKit validity check are included in downstream molecular-property evaluation.


Property Evaluation

For valid generated structures, the molecular properties are recalculated using RDKit and compared with the target conditioning properties.

Four complementary metrics are used.

Mean Absolute Error (MAE)

MAE measures the average absolute difference between generated and target property values.

Lower values indicate smaller average deviations.

Root Mean Squared Error (RMSE)

RMSE measures property error while giving greater weight to larger deviations.

Lower values indicate smaller errors.

Pearson Correlation

Pearson correlation measures the linear relationship between generated and target property values.

Higher positive correlation indicates stronger agreement in the direction of property variation.

Maximum Mean Discrepancy (MMD)

MMD is used to compare generated and reference property distributions.

Lower values indicate closer distributional agreement under the selected kernel.

However, MMD should not be interpreted in isolation.

A compressed generative model can produce a low MMD while generating a narrower or less diverse set of structures whose aggregate properties happen to resemble the reference distribution.

For this reason, MMD is interpreted together with:

  • chemical validity
  • MAE
  • RMSE
  • Pearson correlation
  • qualitative distribution analysis

Distribution Analysis

Selected molecular-property distributions from the final evaluation are shown below.

Molecular Weight

Generated vs Test β€” Molecular Weight

Heavy Atom Count

Generated vs Test β€” Heavy Atom Count

Ring Count

Generated vs Test β€” Ring Count

These plots compare the distributions of valid generated structures with the corresponding reference/test distributions.

They provide a qualitative complement to the numerical evaluation metrics and help reveal differences that may not be captured by a single aggregate metric.


Experimental Pipeline

Open Macromolecular Genome
          Dataset
             β”‚
             β–Ό
      Data Preprocessing
             β”‚
             β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Polymer Representation  β”‚
 β”‚            +            β”‚
 β”‚ 15 Molecular Properties β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚
              β–Ό
hkqiu/PolyTAO-BigSMILES_Version
              β”‚
              β–Ό
 Property-Conditioned Teacher
              β”‚
              β–Ό
     Teacher Hidden States
              β”‚
              β”‚
              β”‚ Feature-Level
              β”‚ Knowledge Distillation
              β–Ό
 Property-Conditioned Student
              β”‚
              β–Ό
       Polymer Generation
              β”‚
              β–Ό
        RDKit Validation
              β”‚
              β–Ό
   Valid Generated Molecules
              β”‚
              β–Ό
  Property Recalculation
              β”‚
              β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Chemical Validity       β”‚
 β”‚ MAE                     β”‚
 β”‚ RMSE                    β”‚
 β”‚ MMD                     β”‚
 β”‚ Pearson Correlation     β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚
              β–Ό
 Capacity–Performance Analysis

Repository Files

The released model includes the standard Transformer checkpoint together with the additional property-conditioning weights required by PolyTAO.

Important files include:

config.json
generation_config.json
model.safetensors
property_proj_student.pt

tokenizer.json
tokenizer_config.json
special_tokens_map.json
added_tokens.json

MolWt_hist.png
HeavyAtomCount_hist.png
RingCount_hist.png

README.md

model.safetensors

Contains the Transformer weights for the released student checkpoint.

property_proj_student.pt

Contains the learned linear projection used to map the 15-dimensional molecular-property vector into the Transformer's hidden space.

This file is required to reproduce the custom property-conditioning mechanism.


Loading the Model

The Transformer and tokenizer can be loaded directly from Hugging Face:

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

MODEL_ID = "FahimehBahman/PolyTAO-Student"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

model = AutoModelForSeq2SeqLM.from_pretrained(
    MODEL_ID
)

model.eval()

print("Encoder layers:", len(model.encoder.block))
print("Decoder layers:", len(model.decoder.block))
print("Hidden dimension:", model.config.d_model)

For the released checkpoint, this should report:

Encoder layers: 2
Decoder layers: 12
Hidden dimension: 768

Loading the Property Projection

The custom property projection is stored separately from the standard Hugging Face Transformer checkpoint.

It can be downloaded and loaded as follows:

import torch
import torch.nn as nn

from huggingface_hub import hf_hub_download

MODEL_ID = "FahimehBahman/PolyTAO-Student"

projection_path = hf_hub_download(
    repo_id=MODEL_ID,
    filename="property_proj_student.pt"
)

property_proj = nn.Linear(
    15,
    model.config.d_model
)

state_dict = torch.load(
    projection_path,
    map_location="cpu"
)

property_proj.load_state_dict(state_dict)
property_proj.eval()

Property-Conditioned Inference

PolyTAO uses a custom property-conditioning mechanism.

Therefore, loading the Transformer checkpoint alone reproduces the Transformer architecture, but does not by itself reproduce property-conditioned generation.

The inference procedure requires:

Polymer Input
      β”‚
      β–Ό
   Tokenizer
      β”‚
      β–Ό
Transformer Encoder
      β”‚
      β–Ό
Encoder Hidden States
      β–²
      β”‚
15 Molecular Properties
      β”‚
      β–Ό
Property Projection
      β”‚
      β–Ό
L2 Normalization
      β”‚
      β–Ό
Scale by 0.05
      β”‚
      β–Ό
Property Embedding
      β”‚
      β–Ό
Conditioned Encoder Representation
      β”‚
      β–Ό
Transformer Decoder
      β”‚
      β–Ό
Generated Polymer

A simplified implementation of the conditioning step is:

import torch
from transformers.modeling_outputs import BaseModelOutput

# properties:
# Tensor of shape [batch_size, 15]
# containing normalized molecular properties.

with torch.no_grad():

    encoder_output = model.encoder(
        input_ids=input_ids,
        attention_mask=attention_mask,
        return_dict=True
    )

    props = properties.clamp(-10.0, 10.0)

    prop_emb = property_proj(props)

    prop_emb = prop_emb / prop_emb.norm(
        dim=-1,
        keepdim=True
    ).clamp(min=1e-6)

    prop_emb = 0.05 * prop_emb.unsqueeze(1)

    conditioned_hidden = (
        encoder_output.last_hidden_state
        + prop_emb
    )

    conditioned_encoder = BaseModelOutput(
        last_hidden_state=conditioned_hidden
    )

The conditioned encoder representation can then be supplied to the model decoder for autoregressive generation.

The property values must be normalized consistently with the preprocessing used during training.


Intended Use

PolyTAO Student is intended primarily for research involving:

  • knowledge distillation
  • Transformer model compression
  • property-conditioned molecular generation
  • polymer generative modeling
  • efficient generative architectures
  • representation-level knowledge transfer
  • capacity–performance analysis

The checkpoint can also serve as a research artifact for studying compressed Transformer models in molecular generation.


Limitations

This model has several important limitations.

  • It is a research prototype, not a production model.
  • Chemical validity does not imply synthesizability.
  • Chemical validity does not imply chemical stability.
  • Generated structures have not been experimentally validated.
  • RDKit validity is a computational validity check rather than experimental verification.
  • Property conditioning is limited to the molecular descriptors used during training.
  • Property values must be normalized consistently with the original training preprocessing.
  • Generation quality can depend substantially on decoding parameters.
  • Distributional similarity does not guarantee structural diversity.
  • A low MMD value alone does not establish superior generative quality.
  • Property metrics are calculated only for structures that pass the validity check.
  • The released checkpoint uses a custom conditioning mechanism outside the standard Hugging Face T5 forward interface.
  • The model should not be used directly for safety-critical, medical, industrial, or experimental chemical decisions without independent expert validation.

Research Context

PolyTAO was developed as part of the M.Sc. thesis:

Knowledge-Distilled Transformers for Property-Conditioned Polymer Generation

The research investigates the effect of Transformer compression on property-conditioned polymer generation.

The work combines:

  • Transformer-based generative modeling
  • molecular-property conditioning
  • feature-level knowledge distillation
  • model compression
  • computational molecular descriptors
  • quantitative evaluation

Multiple student configurations were explored to study how model capacity influences chemical validity, property fidelity, and generated distributions.

The checkpoint published in this repository corresponds to the run configured with:

capacity_percent = 20

and contains:

2 encoder layers
12 decoder layers
d_model = 768

Tools and Libraries

The project uses:

  • Python
  • PyTorch
  • PyTorch Lightning
  • Hugging Face Transformers
  • RDKit
  • NumPy
  • Pandas
  • Scikit-learn
  • Matplotlib

Author

Fahimeh Bahman

M.Sc. Data Science
Software Engineer & Machine Learning Researcher

Research interests include:

  • Generative AI
  • Transformer Models
  • Knowledge Distillation
  • Model Compression
  • Machine Learning Systems
  • AI for Software Engineering
Downloads last month
32
Safetensors
Model size
0.2B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FahimehBahman/PolyTAO-Student

Finetuned
(1)
this model