Classifying fossil teeth with machine learning

My master’s thesis explored how measurements and microscope images can help make sense of tiny dinosaur teeth. This page summarizes the research published with Germán H. Alférez and Keith Snyder in 2025.

A Pectinodon bakkeri tooth with a one-millimeter scale and annotations marking its crown height, basal length, and serrations.
Figure 1 from the publication: the measurements behind the analysis. Select the image to view it full size.

Research context

Teeth can be the only surviving evidence of a dinosaur species. For Pectinodon bakkeri, limited remains make classification especially challenging. Our study used 459 specimens with complete measurements from the Hanson Ranch Bonebed in eastern Wyoming to investigate patterns in their shape.

Method

We examined tooth height, base dimensions, and serrations. Principal component analysis (PCA) helped explain variation in those features; K-means grouped the specimens using the original measurements. Python then organized the corresponding images by group, creating labels for a convolutional neural network built with Keras.

A plot of three tooth clusters: clusters 1 and 3 overlap, while cluster 2 forms a more distinct group below them.
Figure 7 from the publication: three measurement-based groups, with overlap between clusters 1 and 3. Select the image to view it full size.

Clustering results

Further examination indicated that cluster 2 represented another species, though the available remains did not allow a specific identification. We excluded it from image-model training. The two remaining groups overlap; we interpreted this as a possible reflection of differences in tooth size along the jaw.

Building and training the neural network

I built a convolutional neural network in TensorFlow and Keras to classify tooth images into clusters 1 and 3. The model takes 180 × 180-pixel color images as input. Two convolutional layers learn visual features, and max-pooling layers reduce their spatial dimensions. A fully connected layer combines those features before the output layer predicts one of the two groups.

We cropped the images to remove distracting details such as mounting pins and used rotations and horizontal flips to balance the two classes. I normalized pixel values and added dropout layers to reduce overfitting after the initial experiments.

We used 80% of the images for training and the remaining 20% for validation. I trained the network for 100 epochs, using batches of 32 images and the Adam optimizer, and tracked training and validation accuracy and loss to assess how well it generalized.

Image classification results

Our neural network achieved 71% accuracy in distinguishing the two retained groups. Similar tooth shapes, reflections, and inconsistent image backgrounds made that distinction difficult. We proposed separating teeth from their backgrounds and retaining information about physical size as next steps.

My contribution

I developed the software and visualizations and contributed to data curation, analysis, investigation, validation, and the original manuscript alongside my coauthors.