Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device

Xu, Zongzhe

Computer Science > Computation and Language

arXiv:2212.01217 (cs)

[Submitted on 3 Nov 2022]

Title:Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device

Authors:Zongzhe Xu

View PDF

Abstract:This paper proposes a possible method using natural language processing that might assist in the FDA medical device marketing process. Actual device descriptions are taken and matched with the device description in FDA Title 21 of CFR to determine their corresponding device type. Both pre-trained word embeddings such as FastText and large pre-trained sentence embedding models such as sentence transformers are evaluated on their accuracy in characterizing a piece of device description. An experiment is also done to test whether these models can identify the devices wrongly classified in the FDA database. The result shows that sentence transformer with T5 and MPNet and GPT-3 semantic search embedding show high accuracy in identifying the correct classification by narrowing down the correct label to be contained in the first 15 most likely results, as compared to 2585 types of device descriptions that must be manually searched through. On the other hand, all methods demonstrate high accuracy in identifying completely incorrectly labeled devices, but all fail to identify false device classifications that are wrong but closely related to the true label.

Comments:	IEEE Southeast Conference 2023
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2212.01217 [cs.CL]
	(or arXiv:2212.01217v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2212.01217

Submission history

From: Zongzhe Xu [view email]
[v1] Thu, 3 Nov 2022 04:18:05 UTC (330 KB)

Computer Science > Computation and Language

Title:Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators