Open for application · MSc

Enhancing Fine-Grained Perceptual Understanding of Scientific Graphs in Vision Language Models

Supervised by Arlindo L. Oliveira, Hélder Dias

While modern Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general scene understanding, they are frequently hindered by a "global bias" - a tendency to prioritize high-level semantic summaries over local, fine-grained details. As explored in [1], this bias is often a byproduct of pre-training on general image-caption datasets that favor the "gist" of a scene over pixel-level precision. In the context of scientific communication, this limitation becomes a critical failure point; the misinterpretation of a single outlier or a subtle change in slope can lead to entirely inappropriate conclusions. Recent evaluations in MultiChartQA [2] underscore this issue, demonstrating that even frontier models struggle to integrate information across complex, multi-plot figures due to foundational perceptual inaccuracies. Even 37% of all errors of the top-performing model in the study could be attributed to such perceptual issues. Similarly, research in MeasureBench [3] documents systematic fine-grained errors where models consistently misread precise visual markers like axes’ ticks and pointers. This project aims to address this "perception gap" by developing methods that force VLMs to move beyond global heuristics toward a more rigorous, fine-grained interpretation of scientific graphs.

Possible research directions (more ideas welcome)

Domain-specific tokenisation. Current models use a generic "patch-to-token" pipeline that often fragments small details. The student could explore object-centric or coordinate-aware tokenization, where the model learns to represent a graph not as a grid, but as a set of semantically meaningful units. See [4, 5, 6, 7] for more inspiration. Patch-grid alignment & symbolic anchoring. Research in [8] shows that VLM performance drops when a small visual feature (like a data point) shifts from the center of a patch to a boundary. The student could investigate "symbolic anchoring," where the model is fine-tuned to explicitly detect and "mark" axis ticks and legend symbols as a pre-processing step to align the visual grid with the numerical scale. Contrastive locality alignment. Standard VLMs suffer from global bias because they are trained on image-wide captions. This direction explores fine-tuning the vision-language connector using contrastive learning specifically on "local" pairs: for example, contrasting a plot with its correct outlier value vs. a plot with a subtly "shifted" outlier. Neural-Symbolic hybrid interpretation. Building on the DePlot concept [9], this project would move away from simple "image-to-table" translation toward "image-to-coordinate" extraction. The student would develop a module that predicts a sparse numerical representation of the curves/points first, which is then fed into the LLM as a "visual hint," bypassing the noisy perception of the raw pixels during the reasoning phase.

Requisites

The student should have a solid understanding of the Transformer architecture and familiarity with Vision Encoders. They should be comfortable with programming in Python with PyTorch and the HuggingFace ecosystem.

Notes: The selected student will have access to the facilities of INESC-ID and the MLKD group (https://mlkd.idss.inesc-id.pt/), including computing facilities that include four DELL PowerEdge C41402 servers, eight NVIDIA 32GB Tesla V100, four NVIDIA 48GB A40 and four NVIDIA 64GB Tesla A100, among other computing servers (https://mlkd.idss.inesc-id.pt/cluster).

[1] Covert, I., Sun, T., Zou, J., & Hashimoto, T. (2024). Locality alignment improves vision-language models. arXiv preprint arXiv:2410.11087.

[2] Zhu, Z., Jia, M., Zhang, Z., Li, L., & Jiang, M. (2025, April). MultiChartQA: Benchmarking vision-language models on multi-chart problems. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 11341-11359).

[3] Lin, F., Liu, Y., Xu, H., Yue, C., He, Z., Zhao, M., ... & Yang, X. (2025). Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench. arXiv preprint arXiv:2510.26865.

[4] Sun, Z., Ma, Y., Liu, G., Chen, Y., Tang, X., Hu, Y., & Xu, Y. (2026). IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning. arXiv preprint arXiv:2602.03060.

[5] Guo, Z., Diao, E., Yang, C., & Shi, C. (2026). Graph Tokenization for Bridging Graphs and Transformers. arXiv preprint arXiv:2603.11099.

[6] Chen, J., Kong, L., Wei, H., Liu, C., Ge, Z., Zhao, L., ... & Zhang, X. (2024, October). Onechart: Purify the chart structural extraction via one auxiliary token. In Proceedings of the 32nd ACM International Conference on Multimedia (pp. 147-155).

[7] Wu, J., Lu, B., Di, Z., Gan, X., Jin, M., Fu, L., ... & Zhou, C. (2026). : One LLM Token for Explicit Graph Structural Understanding. arXiv preprint arXiv:2602.01771.

[8] Yuan, Y. (2025). Seeing Small: Probing Visual Perception Limits of Vision-Language Models (Master's thesis, University of California, Los Angeles).

[9] Liu, F., Eisenschlos, J., Piccinno, F., Krichene, S., Pang, C., Lee, K., ... & Altun, Y. (2023, July). DePlot: One-shot visual language reasoning by plot-to-table translation. In Findings of the Association for Computational Linguistics: ACL 2023 (pp. 10381-10399).

← All dissertations