Open for application · MSc
Multimodal Document Representation Learning
Learning representations (embeddings) of documents from their constituent elements, such as images, text, and even the layout of the document i.e., the positions of various document components such as text, paragraphs, tables, etc. These embeddings can be used to search for similar documents, which in turn can be used for document classification based on similarity to existing entries in the database, which is something useful for routing documents to different processes.
Concrete implementation suggestions
- Use CLIP [1] with an image encoder and a text encoder, aligning the representations of both. With two distinct encoders sharing an embedding space, it would in principle be possible to search for documents using text and also using images of other documents.
- Follow approach 1 but use a text and layout encoder instead of text alone, or use 3 encoders: one for text, one for layout, and one for images. Using the layout, it would be possible to also search by elements at certain positions in the document.
- Instead of using CLIP, simply use the encoder of a multimodal model already pre-trained on documents (UDOP [2], LayoutLMv3 [3], etc.) with the various modalities described above and train the embeddings to bring similar documents closer together. This approach has the added difficulty of needing to aggregate similar documents and define what makes a document similar, making it not 100% self-supervised.
Cooperation with external entity: Fidelidade S.A.
Requisites
We value strong proximity of students with Fidelidade, so it would be ideal if students could be present at the office on at least some of the team's in-person working days.
- [1] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., … Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision . arXiv [Cs.CV]. Retrieved from http://arxiv.org/abs/2103.00020
- [2] Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., … Bansal, M. (2023). Unifying Vision, Text, and Layout for Universal Document Processing . arXiv [Cs.CV]. Retrieved from http://arxiv.org/abs/2212.02623
- [3] Huang, Y., Lv, T., Cui, L., Lu, Y., & Wei, F. (2022). LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking . arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2204.08387