Computing embeddings and dimensionality reduction#

In this notebook we load a list of PhD topics and create an LLM-embedding from them. In such an embedding, each PhD topic is represented in high-dimensional space, e.g. as a vector with 1000 numbers. In order to display these embeddings on screen, e.g. in a two-dimensional plot, we apply dimensionality reduction to it.

from openai import OpenAI
import pandas as pd
from sklearn.manifold import TSNE
from umap import UMAP
import stackview
import numpy as np
import yaml
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

First, we load the CSV file and take a look at it.

df = pd.read_csv("phd_topics.csv")
df
name research_field topic
0 Taylor Reed Software Engineering Runtime Verification of Adaptive Microservice ...
1 Riley Jain Robotics Learning‑Based Resilient Coordination of Heter...
2 Taylor Adams Computer Vision Unsupervised Spatiotemporal Representation Lea...
3 Devon Thomas Machine Learning Uncertainty‑Aware Reinforcement Learning for A...
4 Alex Lee Computational Complexity Fine‑Grained Complexity of Dynamic Subgraph Is...
... ... ... ...
245 Kris O'Hara Computer Vision Self-Supervised Transformers for 3D Scene Reco...
246 Riley Thomas Human-Computer Interaction Emotion‑Adaptive Conversational Interfaces via...
247 Skyler Jain Human-Computer Interaction Adaptive Multimodal Interaction Models for Emo...
248 Dana Brooks Artificial Intelligence Hybrid Neural–Symbolic Architectures for End‑t...
249 Morgan Singh Programming Languages Design and Verification of an Algebraic Effect...

250 rows × 3 columns

Second, we load the embedding model intfloat/multilingual-e5-large-instruct, a leading small embedding model.

e = HuggingFaceEmbedding(model_name="intfloat/multilingual-e5-large-instruct")
e.get_text_embedding("Hello world")[:5]
[0.00521087646484375,
 0.0252685546875,
 0.007358551025390625,
 -0.044891357421875,
 0.02490234375]

Next, we test this model.

vector = e.get_text_embedding("Hello world")
len(vector), vector[:5]                              
(1024,
 [0.00521087646484375,
  0.0252685546875,
  0.007358551025390625,
  -0.044891357421875,
  0.02490234375])

The following code will apply the embed function to all topics in our table.

df["embedding"] = df["topic"].apply(e.get_text_embedding)
df.head()
name research_field topic embedding
0 Taylor Reed Software Engineering Runtime Verification of Adaptive Microservice ... [-0.0015439987182617188, 0.027496337890625, -0...
1 Riley Jain Robotics Learning‑Based Resilient Coordination of Heter... [0.0090789794921875, 0.0286865234375, -0.01675...
2 Taylor Adams Computer Vision Unsupervised Spatiotemporal Representation Lea... [-0.00688934326171875, 0.01209259033203125, -0...
3 Devon Thomas Machine Learning Uncertainty‑Aware Reinforcement Learning for A... [-0.0050506591796875, 0.0226898193359375, -0.0...
4 Alex Lee Computational Complexity Fine‑Grained Complexity of Dynamic Subgraph Is... [0.03546142578125, 0.03033447265625, -0.026321...

Again, we apply dimensionality reduction for visualization purposes, namely t-SNE and UMAP.

# Convert embedding vectors to numpy array for t-SNE
embeddings = np.array(df['embedding'].tolist())

# Apply t-SNE
tsne = TSNE(n_components=2, random_state=42)
tsne_embeddings = tsne.fit_transform(embeddings)

df['TSNE0'] = tsne_embeddings[:, 0]
df['TSNE1'] = tsne_embeddings[:, 1]
# Convert embedding vectors to numpy array
embeddings = np.array(df['embedding'].tolist())

# Apply UMAP
umap = UMAP(n_components=2, random_state=42)
umap_embeddings = umap.fit_transform(embeddings)

df['UMAP0'] = umap_embeddings[:, 0]
df['UMAP1'] = umap_embeddings[:, 1]
C:\Users\rober\miniforge3\envs\rag26\Lib\site-packages\umap\umap_.py:1952: UserWarning: n_jobs value 1 overridden to 1 by setting random_state. Use no seed for parallelism.
  warn(
df["selection"] = 1

The resulting two dimensions can be visualized on screen.

stackview.scatterplot(df, column_x="UMAP0", column_y="UMAP1")
df["selection"].unique()
array([1])

Finally, we store the topcis, together with the embeddings and the two-dimensional UMAPs to a yml file.

# Convert DataFrame to dictionary
data_dict = df.to_dict()

# Save as YAML file
with open('phd_topics.yml', 'w') as file:
    yaml.dump(data_dict, file)