Code for running SPMat algorithm Please cite it as :
In this work, we propose supervised pretraining for material property prediction, where available class information serves as surrogate labels to guide learning, even when downstream tasks involve unrelated material properties. We evaluate this strategy on two state-of-the-art SSL models and introduce a novel framework for supervised pretraining.
The CGCNN model we pretrained using SSL has been tested on the databases from Materials Project. We pretrained the SSL models with 121k crystal structures and their class labels from Material Project and tested the performance after finetuning for six different regression tasks on 33990 structures from Material Project.
To run the SPMat code the following packages are required
It is advised to create a new conda environment and then install these packages. To create a new environment please refer to the conda documentation on managing environments (https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html)
To input crystal structures to SPMat, you will need to define a customized dataset. Note that this is required for both training and predicting.
The dataset that we use for this work are in the cif format.
- CIF files recording the structure of the crystals that you are interested in
- The values of the target properties for each crystal in the datase
You can create a customized dataset by creating a directory root_dir with the following files:
-
atom_init.json: a JSON file that stores the initialization vector for each element. Theatom_init.jsonfile has some of the basic atomic features encoded. Please refer the supplementary information of the paper to find out more about the basic atomic features. -
ID.cif: a CIF file that recodes the crystal structure, whereIDis the uniqueIDfor the crystal
The models have been pretrained on the cif files in the Material Project database. In total we aggregate 121k cif files for the pretraining. While pretraining, we used class labels available in Material Project to guide the process effectively making it a supervised pretraining in SSL domain. To run pretraining you will need id_prop.csv inside datasets folder containing the name of the cif file and its class information (metal/nonmetal, conductor/semiconductor/insulator). Create one additional folder in the root named runs_contrast to save the pretrained models. Also don't forget to keep your pretrining .cif files inside pre_train_mp folder. Currently it does not have any .cif file. To run the pretrained model run the command
python contrast.pyThe parameters for the pretraining of model can be modified in the config.yaml and config_simclr.yaml
For Finetuning the model, we initialize with the pre-trained weights and finetune it for the downstream task. You need the pretrained models saved in runs_contrast to get pre-trained weights and also finetuning .cif files saved in the designated folders. Please see the config_finetune_new.yaml file to see how to name the folders to keep finetuning cif files. For your own dataset, you need to have id_prop.csv in the folders containing the finetune cif files. Create an additional folder named runs_ft in the root to store the finetuned models. Also, create a folder named 'experiments' to save the finetune results as csv files.
python finetune_cgcnn_spmat.pyFor datasets in the Materials Project - MP.
- CGCNN: https://github.com/txie-93/cgcnn
- Barlow Twins https://github.com/facebookresearch/barlowtwins
- Crystal Twins: https://github.com/RishikeshMagar/Crystal-Twins
- SupCon: https://github.com/HobbitLong/SupContrast
The paper titled by "Supervised Pretraining for Material Property Prediction" is uploaded to Arxiv: https://arxiv.org/abs/2504.20112
