The research let us fine-tune the model into an analysis engine that can tell whether texts were generated by the same language model or written by the same person.
Research uncertainty: hypotheses were tested iteratively with no guaranteed outcome. A limited compute budget for fine-tuning.
We carried out applied research on multilingual-e5-base: we studied the language model’s sensitivity to writing styles and proposed a fine-tuning method that reduces errors in cross-lingual tasks. The research produced a Gabriel Graph that shows which authors wrote similar texts. If the model judged two texts “mutually sensitive,” the graph linked their authors. The results have been integrated into several of the client’s language models and serve as a “temperature” indicator of how interconnected the data is.
- 0101
Problem research
Reviewing existing approaches to text style extraction and metric-learning architectures.
- 0202
Dataset preparation
Assembling a text corpus and labeling authorship and style features.
- 0303
Neural network architecture
Building models with ArcFace, SphereFace, and CosFace loss functions.
- 0404
Contrastive learning
Training the system to extract stylistic features of text with contrastive approaches.
- 0505
Quality evaluation
Measuring author identification accuracy on short text fragments.
- 0606
Research write-up
Compiling the results and extending the work toward detecting AI-generated text.
- 01PyTorch
- 02HuggingFace
- 03multilingual-e5-base
- 04Python
- 05Streamlit
- 06CUDA
- 07TensorBoard
- 08LangChain
Have a similar challenge?
We’ll review it and estimate within 5 business days.
