Leak It: Per-Document Extraction Beyond Aggregate Membership Inference
arXiv preprint, 2026
A probabilistic method for extracting and quantifying memorized training data from black-box language models without white-box access to model internals.
research
Research across language models, computational biology, and health. The latest citation record is on Google Scholar.
arXiv preprint, 2026
A probabilistic method for extracting and quantifying memorized training data from black-box language models without white-box access to model internals.
bioRxiv preprint, 2026
Benchmark dataset of 112,399 missense variants across 2,809 genes, each labeled as GOF, LOF, or neutral by board-certified clinical geneticists. Public on Kaggle and Hugging Face.
SSRN preprint; poster accepted, XLIII Congresso Brasileiro de Psiquiatria (CBP) 2026, 2026
A nationwide registry-based re-estimation of Indigenous suicide mortality in Brazil through 2024, using the complete national mortality registry and the 2022 Census to characterize its age, sex, regional, and method structure.
EMNLP (Industry Track), 2025
A benchmark and methodology for evaluating small function-calling language models on realistic e-commerce dispatch tasks under tight latency budgets.
Simpósio Brasileiro de Sistemas de Informação (SBSI), 2025
ML approach to predict whether a missense variant produces a gain- or loss-of-function effect, with applications in clinical variant interpretation.
Information, 2024
A bibliometric study mapping the GNN research landscape across applications and sub-fields.
Frontiers in Chemistry, 2022
Review of graph neural network architectures applied to virtual screening pipelines for drug discovery.
Analytical Methods, 2020
PLS-regression model on ATR-FTIR spectra to quantify schizophyllan directly in fermented broth, avoiding costly downstream isolation.