Introduction
Certara.AI Curate assists medical writers and scientists in the creation of reports and other documents via an easy-to-use web application.
Requirements
Curate is supported on the following browsers:
· Chrome
· Safari
· Edge
· Firefox
Getting Started
Once you are added as a user to Curate, you will receive an email with instructions on how to access the application.
Registration Link Steps
If you are sent a registration link, click on the link to navigate to the Registration Page.
1. Enter your Full Name
2. Enter a Username
3. Enter your E-mail address
4. Create a Password that is 8+ characters long
5. Click the box next to “I agree to the Terms of Service”
6. Click on the “Register” button
Reset Password Steps
You may receive an email with a username and a temporary password. Click on the link to navigate to the Sign In page.
1. Enter your Username and temporary password to sign in
2. Reset your password by clicking on “Forgot password?”
a. Enter your email address
b. Click “Reset”
You will receive an e-mail with instructions to reset your password.
What is Certara.AI Curate?
Certara.AI Curate helps scientific organizations build, optimize, and maintain AI models to analyze and extract meaningful information from unstructured data.
What’s the problem with unstructured data?
Approximately 80% of a given company’s data is “unstructured” - a.k.a. documents, slides, posters, etc. - because those formats make it easier for people to process and extract meaningful information. However, computers like “structured” data - a.k.a. spreadsheets, databases, and so on. It’s incredibly labor-intensive to apply any kind of data analytics to unstructured content since computers need a hand to extract the information we implicitly understand. This leaves organizations with a massive hole in what they know!
Here’s how Certara.AI Curate helps:
AI helps by learning what types of information people want to extract from their massive piles of unstructured data
People can extract this information by providing a set of documents to Curate’s AI tools and asking questions, like “What ADC manufacturers are listed in this paper?”
The AI model “reads” these documents to find answers to the question, returning what it thinks is the best result and where it found it in context
Curate arranges these results in a smart spreadsheet
People can verify these answers, provide feedback to the model (so it’ll do better next time!), and export their results either within the app or via an API to power further analytics
Certara.AI Curate provides structure to unstructured content.
Understanding the Certara.AI Curate Layout
The user interface (UI) of Certara.AI Curate comprises two distinct sections, namely, the top toolbar and the main view. The top toolbar facilitates seamless navigation to different applications and buttons within the UI while the main view displays the content per the selected button.
Application Menu
The menu of applications provides users with a seamless experience in navigating various applications.
Document Sets
The button labeled "Document Sets" presents a comprehensive view of all existing documents and their respective availability. Additionally, it provides the option to create a new Document Set and establish a separate folder by utilizing the Actions Button.
Create Document Set
Select "Create Document Set" to create a repository to organize documents by keywords, dates, and other parameters. This approach serves to minimize the number of documents that require review and enables prompt access to specific sets of documents when necessary.
Create Folder
"Create a folder" to organize document sets based on keywords, creation dates, and other relevant filters. This will reduce document clutter and help you access relevant sets that must be reviewed in response to a specific query and enhance the efficiency of your document management processes.
Batches
The "Batches" section is the one-stop-shop for a complete overview of all created batches. It comprises essential information such as Batch Key, Question Count, Progress, Type, and Date Created. By utilizing the "Actions" Button, you can confidently create custom QA or Ontology Batches that align with your requirements.
Create QA Batch
"Create a QA Batch" enables the creation of a batch of questions where you can further filter by keywords, page range, document set, document sections, and deep learning model.
Create Ontology Batch
The "Create Ontology Batch" enables the creation of a batch of Ontology QA enabling the review and ensuring the utilization of a vector database populated by an ontology of harmonized terms.
My Assignments
"My assignments" is designated to provide you with a comprehensive list of items that have been assigned for your review. This section serves as a central repository for all assignments designated to you.
Models
Within the "Models" section, you will find a comprehensive list of both "Classification Models" and "QA Fine-Tuning Projects". Additionally, you have the option to create your own Classification Model or Fine-Tuning Project by utilizing the "Actions" Button.
Create Classification Model
"Create Classification Model" empowers you with the ability to build out a ground truth training dataset for the classification model to train on. You can use many filters as seen in Document Sets to assist you in this process.
Create QA Fine-Tuning Project
"Create QA Fine-Tuning Project" enables you to create designated project folders that serve as a repository for various Models and iterations, specifically tailored to a particular use case and all housed within a singular folder.
Building and Maintaining a Document Set in Curate
Document sets enable you to bundle multiple related documents together so that you an organize and manage related documents as a single entity. Let’s explore how to build and maintaining document sets!
Building a Document Set
First, create a new document set.
1. Open “Curate”
2. Select “Create Document Set” from the “Actions” menu.
3. Use the filters on the right-hand sidebar to narrow the document set to the documents you would like to interrogate with the Large Language Model.
a. Alternatively, you can drag and drop documents onto the right-hand sidebar to automatically add them to the document set.
4. To find the documents you are looking for, simply enter the relevant keyword(s) in the search field under the SEARCH tab. For more information on how to refine your search, check out our article Boolean Operators in Layar.
5. Once the set contains the documents you would like to work with, Click the “Save” button to Save the document set.
6. After saving, select “Ask Question” from the “Actions” menu.
7. In the modal window that appears, type the question you would like answered.
8. Click on “Show Advanced Settings” and
9. Select “Large Language Model” from the “DEEP LEARNING MODEL” drop-down
Curate will begin a batch process and will start answering your question for each document in the document set. A new column will appear in your document set with the responses from the Large Language Model.
Adding Documents to a Document Set
Add documents to a document set with the following methods.
Drag and Drop
Log into Certara.AI Curate
Click to select a document set or create a new one
Drag and drop documents into your browser onto the right-hand sidebar or click the link in the right-hand sidebar
Upload Documents
1. Log into Certara.AI Curate
2. Click to select a document set or create a new one
3. Select “Upload Documents” from the “Actions” menu
4. Navigate to the document
5. Click to select the document
6. Click on the Open button
How to Create a Document Set in Curate
To organize related documents by keywords, dates, and other parameters, you can create a Document Set. This approach helps to minimize the number of documents that require review and enables quick access to specific sets of documents when needed. Let’s walk through the steps:
1. Log into the "Curate User Interface (UI)"
2. Click the "Actions" button and
3. Select "Create Document Set"
You will see a brand-new Document Set, with no filters, other metadata columns, or questions asked of the document. This has every document in your data fabric!
4. To filter and display only the relevant documents you wish to add to your Document Set, follow these steps:
a. Enter one or more keywords in the free text field labeled as "SEARCH" (Check out the Boolean tips page to learn more about conducting searches for narrowing the document set.)
b. Click on the enter or return key to start the search
c. Utilize the available filters to further refine your search and narrow down the results.
5. Click "Save"
You will see a popup at the bottom right corner of your screen indicating that the document set was successfully saved.
How to Create a Folder in Curate
Create a Folder to organize document sets based on keywords, creation dates, and other relevant filters. This will reduce document clutter and help you access relevant sets that must be reviewed in response to a specific query and enhance the efficiency of your document management processes.
Ensure that you have a document set or sets that you want to add to the folder before proceeding with the steps to create a folder. Create a document set by following the instructions provided in the article "How to Create a Document Set in Curate".
To Create a Folder, complete the following steps:
1. Log into the "Curate User Interface (UI)"
2. Click the "Actions" button
3. Select "Create Folder"
4. In the "Create Folder" modal, enter a name in the free-text field under "NAME"
5. Click "Save"
Your newly created folder will be listed in the Document Sets table and the document sets. You will see a folder in the second column indicating it is a folder.
6. Click on the box next to the document set or sets you would like to add to the folder
7. Click on the "Selected" button
8. Click on "Move"
9. Click on the "New Folder"
10. Click on the "Choose Destination Button"
🛑 Pro Tip: Folders will disappear if no documents are added upon creation.
Asking Natural Language Questions in Curate
Using Curate you can use the Large Language Model capability of Composer to ask questions of documents in a Curate document set. A license for Composer is required for this functionality to be enabled.
First, create and save a new document set. Instructions can be found at the following link:
Ask Questions
Once you create a document set, you can Ask a Question from the Actions menu.
In the modal that appears, type the question you would like answered.
1. Click the “Show Advanced Settings” link
In the Ask Question model, under the DEEP LEARNING section, there are two options: BioMegaTron SQuAD v2 and Large Language Model
BioMegaTron SQuAD v2 is an extractive model: it reads through documents to find the exact line of text that includes the answer to your question
The Large Language Model is abstractive: it reads through documents and provides either a summary of the text to answer your question or infers the answer from the text
You will also find the custom models you’ve retrained and published here.
💡 Pro Tip - Before attempting to apply a newly developed query on a larger set of documents, it is recommended to perform prompt engineering on a smaller subset of around 20 documents in Curate.
Model Retraining Project
A Question-Answering (QA) project can be perceived as a designated folder that serves as a repository for models, along with their iterations, that are tailor-made for a specific use case. The folder is intended to streamline the management and accessibility of the models, while also ensuring that all the relevant iterations and supporting material are stored in one place.
You can find these projects under the “Models” tab in the Curate UI.
QA Model Parameters
Base Model - The pre-trained QA model is used as a starting point for fine-tuning. Users have the option of starting with BioBert, Distilbert, or RoBerta large language question answering (QA) models. If there is another model that is of interest to leverage as your initial base model, please reach out to your Certara.AI contact.
Epochs - An epoch is the number of times a model sees a phrase or an example input from the training corpora.
Test Ratio - This is the ratio of testing/training data. Of the documents you assigned to each class, a portion of each dataset will be separated from the training data and used to test the model afterward. By default, the test set size is set to 0.1, or 10% of your input sets.
Batch Size - Number of examples in a single mini-batch optimization step, used as part of the learning rate optimization.
Max Length - Maximum length of question+context pairs. For shorter documents, like tweets or PubMed abstracts, this can be reduced. Smaller numbers will mean faster training and inference.
Early Stop - Prevents overtraining/overfitting of the model by stopping training if certain criteria are met.
QA Model Performance Metrics
Exact Match (EM) - Exact Match is a common metric combined with an F1 score to determine a model’s performance. Exact match means the answer predicted by the model is an exact string match to the ground truth training dataset. If the model prediction for a document is an exact match to the ground truth, the value is 1. The exact match metric is therefore an aggregate of all of the model’s predictions that were found to be an exact match. These include moments where the model correctly predicted a text answer, as well as when the model correctly predicted there was no applicable answer (see EM No Answer and EM Text Answer for the splits into each distinct group).
Note: Since exact match is binary, and F1 scores are more flexible, it is almost always the case where exact match scores will be lower than the F1 score. This being said, a high F1 score can still strongly indicate a highly performant model. There’s no easy choice between using EM or F1 scores to determine the model’s improvement, so we suggest using both as part of your model performance evaluation.
EM No Answer - Refers to the number of times where the model predicted “no answer applicable”, which matched then the ground truth training dataset also determined there was no answer applicable.
Note: We often refer to these as “true negatives”, and they are critical to fine-tuning a model to make sure the recall is as high as possible for the QA model, without sacrificing the accuracy of the answers provided (after all, no one wants a model to return an answer just for the sake of returning something). It is very common to have a very low number of true negatives, but they are invaluable to improving the accuracy and recall of your model.
EM Text Answer - Refers to the number of times where the model was able to provide a prediction derived from the text that is an exact match to what was found in the ground truth dataset.
Note: These are often referred to as “true positives”, and are the most common type of class generated for both the ground truth dataset.
Total No Answer - A total number of correctly predicted “no answers” based on the demographics of the test set.
Total Text Answer - Total number of correctly predicted text-referenced answers based on demographics of the test set.
F1 Score - Compared to an exact match, the F1 score metric is a more granular metric that measures the average overlap between the prediction and ground truth answer. The metric balances the goals of precision and recall to determine whether the model was able to get to the sentence where the answer resides. For example, take the sentence: “My doctor prescribed me 1000 mg acetaminophen to take after my surgery.” The ground truth answer may be “acetaminophen”. An exact match will only score the model as correct if it predicts “acetaminophen” exactly, whereas an F1 score will be more lenient and say the model was 79.3% accurate* when it guessed “1000 mg acetaminophen”. It’s correct, just not exactly the answer the ground truth provided. *The calculation was arbitrarily made for the sake of the definition.
F1 Text Answer - This is the F1 score specifically for text-derived answers. The “no answer applicable” true negatives are excluded from this computation.
Top N - Oftentimes, the QA model will return several (n) number of answers for a given document. This parameter details how many of the top answers (ranked by probability score) were included in the performance metrics detailed below. For example, Top N = 4 means that there was a correct answer within the top 4 answers provided by the model for a given document.
Top N EM Text Answer - Refers to the number of times where the model was able to provide a prediction derived from the text that is an exact match to what was found in the ground truth dataset, within the top n (e.g. top 4) answers predicted by the model.
Top N Accuracy - For QA accuracy, an answer is given either a 1 or 0. If a predicted answer has any overlap with the label/correct answer, it will be given a score of 1. For example: A predicted answer of "San Francisco" when the answer is "San Francisco, California" would receive a score of 1. A predicted answer of "Los Angeles" would be 0. So if we assume for Top 1 Accuracy for 100 examples, 90 predictions had appropriate overlap and 10 did not, the accuracy would be 90%.
Note: The accuracy metric is intended for closed systems tasks (in this case, the overlap is happening on the specific instance of the answer). So, if the answer is "San Francisco," and it appears twice in a document, it needs to overlap with the correct instance of San Francisco (whether that be the first or second), that was defined in the ground truth dataset. For EM and F1 scores, it doesn’t matter whether the answer was found in the first or second instance, just that the answer’s string (San Francisco) was correct.
Top N Accuracy Text Answer - This is the same accuracy metric as the Top N Accuracy, but the “no answer applicable” true negatives are not included in the computation.
Top N F1 Text Answer - This is the F1 score specifically for text-derived answers, where the correct answer was found within the top n (e.g. top 4) answers predicted by the model for a given document. The “no answer applicable” true negatives are excluded from this computation.
How to Publish / Unpublish a QA Model
Once you are satisfied with your model, you can publish it to be used as the model for QA in Curate!
How to Find Your Model in the Model View
1. Log into “Certara.AI Curate”
2. Click on “Models” in the top toolbar
On the Models page, you will see the Classification Models Table with the Name, State, and Date Created.
3. Click to select a Name load and work with the Model.
How to Publish a QA Model
1. Click on the model you wish to publish
2. Click the “Actions” button and
3. Select “Publish”
How to Unpublish a QA Model
1. Click the “Actions” button
2. Select the “Unpublish” option, which is only visible to published models
How to Create a New QA Fine-tuning Project
QA Projects are designated project folders that serve as a repository for various Models and iterations, specifically tailored to a particular use case and all housed within a singular folder.
The Models View
You can get to the Models view by clicking on the "Models" menu option, located in the top navigation bar of the Curate application.
This is the default view with no projects listed in the QA Fine-Tuning Projects section. Create a new project to start creating a model.
The current view displays an absence of any projects within the QA Fine-Tuning Projects section. To commence the process of developing a model, create a new project.
Create QA Fine-Tuning Project
1. In your Models view, click the Actions button and
2. Select “Create QA Fine-Tuning Project
Note: Retraining only uses human-validated answers from retraining set columns. At least 50 validated answers per column are recommended for a meaningful increase in performance.
3. In the “Create QA Fine-Tuning Project” popup, enter a name for your project, then click Save
For example, if you are building a QA model that will yield answers around diseases and symptoms, you might name the project “Diseases”.
Once you create a QA Fine-Tuning Project, you will see it listed in the QA Project's dashboard view.
You can track models here as you build them.
Create Model
Learn how to create a new QA fine-tuning model within your QA project.
1. From the Models View, click on the name of the project listed under “QA Fine-Tuning Projects”
2. Once you are in the Project's working view, click on the “Actions” button
3. Select “Create Model”
You will be guided through defining various parameters to build a QA model.
Step 1: Select Questions Columns
4. Click the “Add Question Columns” button
This will open a “Select Questions Column” modal where you can select questions from different Document Sets to include in your training dataset.
5. Select questions from the dropdown menu or use the search bar to find specific ones
6. Click “Done” when you are finished with your selection(s)
After selecting the training inputs for your ground truth training dataset, move to the custom parameters phase of the model build job to configure them.
Step 2: Customize
Once you have solidified your training dataset, you now have the option to adjust any model parameters used during the model fine-tuning.
We have already provided default inputs should you wish to skip this step and use the base parameters instead. See QA Model Parameters for descriptions of each advanced parameter option.
Step 3: Train Model
Once you have determined your parameters for the QA model, you can now click the “Train Model” button in the third panel.
Upon clicking "Train Model", you’ll be asked to save a name for your new model. If you plan to do multiple iterations of the model, we suggest having some sort of versioning in your name (e.g. “Disease QA Model V1”, “Disease QA Model V2”, etc.).
💡Pro Tip - The name can be as vague or verbose as you please, though we often see our customers naming these based on a combination of things:
1. The model’s name (if you’ve coined it)
2. The dataset you curated for training (e.g. “MS-ClinicalTrials”)
3. The date the training had commenced
4. The iteration (version) of the model
A full name might look something like “ClinBert_MS-ClinicalTrials_Sept2024_V1”. However, it’s worth noting that all of these inputs are reference-able using the Curate Model dashboard, so this is purely a preferential choice.
Upon clicking "Save", your model will begin automatically building, using the training datasets and parameters you have defined. You can leave or close out of this page and return any time by clicking on the “Models” tab and into your “QA Project” of interest.
When the model is done building, this page will automatically update to show you the model’s performance results, which were run on the test set you determined in your parameters portion of the training (see: Test Ratio in QA Model Parameters).
To learn more about these performance metrics, please see QA Model Performance Metrics.
How to Retrain a QA Model
If you would like to retrain a previously built model, either by adding new training data or by tinkering with custom parameters (or both!).
Before you begin, navigate to your Model in the Model View Tab.
1. Click on the model you’d like to use as your base model
2. In the Actions button located in the top right corner, click the Retrain Model option
3. Add question columns to adjust the dataset or parameters. Only human-validated answers from those columns will be used during retraining (a minimum of 50 validated answers per column is recommended for meaningful performance improvement).
Here is an example of where the EPOCHS was increased from 3 to 6.
See QA Model Parameters for descriptions of each advanced parameter option.
4. Click “Train Model”
5. Type a name in the free-text field under “Name” in the “Save Modal” pop-up
Compare the results between this model’s performance metrics and the prior model’s metrics! Continue to tinker until you are satisfied with your model’s performance, at which point you can publish your model for use within the software (see How to Publish / Unpublish a QA Model).
How to Create a QA Batch
"Create a QA Batch" enables the creation of a batch of questions where you can further filter by keywords, page range, document set, document sections, and deep learning model. Complete the following steps to Create a QA Batch!
1. Log into the "Curate"
2. Click on "Batches"
3. Click on the "Actions" Button
4. Click on "Create QA Batch"
5. In the "Create Batch" modal, enter a "BATCH KEY" and "QUESTION KEY"
6. Enter "QUESTION STRINGS"
7. Add optional "KEYWORDS"
8. Enter optional "PAGE RANGE"
9. Select a "DOCUMENT SET" from the menu
10. Select optional "DOCUMENT SECTIONS"
11. Select "DEEP LEARNING MODEL" from the menu
12. Click "Create"
Classification Model Parameters
We utilize fastText as our base classification model, which you can learn more about from their documentation here.
Test Set Size (Integer) - This is the ratio of testing/training data. Of the documents you assigned to each class, a portion of each dataset will be separated from the training data and used to test the model afterward. The test set size is set to 0.1, or 10% of your input sets by default.
Rebalance - This feature rebalances the input sets so that there is a more equal distribution of labeled data within each class.
This does result in an unequal testing set distribution, as the test set is separated from the training dataset before rebalancing is applied.
Learning Rate - The learning rate of an algorithm indicates how much the model changes after each example sentence is processed. We can both increase and decrease the learning rate of an algorithm. A learning rate of 0 means that there is no change in learning, or the rate of change is just 0, so the model doesn’t change at all. The usual learning rate is 0.1 to 1.
Learning rates can also be optimized automatically. There are several available optimizers publicly available for learning rate optimization. We currently use Adam as our optimizer, an adaptive learning rate optimizer specifically trained for deep neural networks.
Dimensions - The pre-trained word vectors we distribute have a default dimension size, or dimensionality, of 100. This default parameter is different from the default used by the standard fastText base model (which uses 300 dimensions). To increase or decrease the vector dimensionality, adjust this parameter accordingly.
What is dimensionality?
Dimensionality is the size of the word vectors when they are abstracted into what we refer to as “vector space”, which consists of many hundreds or thousands of planes of dimensionality. Dimensionality Reduction (DR) algorithms transform a set of N high-dimensional input data points with M dimensions into an output dataset with the same number of data points but a reduced number of m dimensions: m < M. This technique is used to generate tractable input datasets for classification models, and also aids in visual data analysis application, such as the build-out of t-SNE, UMAP, and PCA plots.
Window Size (Int, Optional) - The maximum distance between a sentence's current and predicted word. This is also the size of the context window.
Epochs - An epoch is the number of times a model sees a phrase or an example input from the training corpora.
Min Word Count (Int, Optional) - The model ignores all words with a total frequency lower than this value.
Min Label Count - The model ignores all labels with a total frequency lower than this value.
This is particularly important if you are using a multi-class label, where your training dataset may only have a handful of examples for a given label. If this is the case, you’ll likely want to consider either (1) gathering more examples for those labels to train on or (2) removing this label from your model.
Min Character NGram Length (Int, Optional) - Minimum length of character n-grams to be used for training word representations.
Max Character NGram Length (Int, Optional) - Maximum length of character ngrams to be used for training word representations. Set max ngram length to be lesser than min ngram length to avoid character ngrams being used.
Max Word NGram Length (Int, Optional) - Maximum length of word n-grams to be used for unigram, bigram, trigram, etc. word handling.
Negatives Sampled (Int, Optional) - If > 0, negative sampling will be used, where the int value for this parameter specifies how many “noise words” should be drawn (usually between 5 and 20). If set to 0, no negative sampling is used.
Loss Function - Loss functions play an important role in any model. There are too many unknowns when calculating the perfect weights for a neural network, which is why the problem of learning is cast as an optimization problem and we utilize loss function algorithms to navigate the space of possible sets of weights the model may use to make good or “good enough” predictions.
Loss functions define an objective which the performance of the model is evaluated against and the parameters learned by the model are determined by minimizing a chosen loss function. In short, loss functions define what a good prediction is and isn’t.
By default, we use the softmax algorithm as our loss function. However, users can also select from three additional algorithms as they choose. We will provide additional documentation on these algorithms by request.*
Available Loss Functions
Softmax (default)
NS - Skipgram negative samples, or SGNS
HS - Skipgram hierarchical softmax
OvA - “One-vs-All”, also referred to as “One-vs-Rest” or OvR
If interested in using other loss functions not currently provided through Signal Curate, we suggest reaching out to your Certara.AI support team for additional guidance.
Buckets (Int, Optional) - Character ngrams are hashed into a fixed number of buckets, to limit the memory usage of the model. This option specifies the number of buckets used by the model. The default value of 2000000 consumes as much memory as having 2000000 more in-vocabulary words in your model.
Learning Rate Update Rate - As mentioned earlier, learning rates can be optimized automatically. There are several available optimizers publicly available for learning rate optimization. We currently use Adam as our optimizer, an adaptive learning rate optimizer that utilizes what’s known as an update rule to speed up optimization performance in terms of speed of training. This update rule is derived from model weights and step size.
If you plan to tinker with this update rate, we suggest connecting with Certara.AI's data scientists to best evaluate different learning rate optimizers for your specific use case.
Sampling Threshold - The threshold for configuring which higher-frequency words are randomly down-sampled, a useful range is (0, 1e-5).
Classification Model Performance Metrics
Precision, Recall & F1 Scores
Precision is the number of true positives divided by the number of true and false positives. Precision can be thought of as a measure of a classification model’s exactness. A low precision can also indicate a large number of false positives.
For example, imagine we have two oranges and two apples, and we have trained a model to classify fruits as either orange or apples. If the model predicts each of the four correctly, the model gets a 100 on precision. If the model instead predicts the two oranges correctly but mistakes one of the apples for an orange, the precision would be 75% (3 true positives, 1 false positive).
Now, imagine we have two oranges, an apple, and a pear. If the model predicts the oranges and apples correctly and does not provide a label for the pear, the precision would still be 100. Since there’s no class for the pear, not returning a label is considered a true negative (and thus not included in the calculation) so the calculation would be 3 true positives divided by (3 true positives + 0 false positives).
Recall is the true positives overall predicted results (true positives, false positives, true negatives, and false negatives). Put another way, recall is a measure of a classification model’s completeness. A low recall can indicate a large number of false negatives.
Take the example from before: we have trained a model to classify fruits as either orange or apples. First, we have two oranges and two apples, and the model guesses all four correctly. The recall is going to be 100. If the model instead predicts the two oranges correctly but mistakes one of the apples for an orange, the recall would be 75% (3 true positives, 1 false positive). If the model instead predicts the two oranges correctly, predicts one apple as an orange, and doesn’t provide a label for the final apple, the recall would be 50% (2 true positives, 1 false positive, and 1 false negative).
Let’s take it one step further: we have two oranges, an apple, and a pear. If the model predicts the oranges and apples correctly and does not provide a label for the pear, the recall would be 100 (3 true positives, 1 true negative). Since there’s no class for the pear, not returning a label would be considered a true negative.*
*True negatives are a coveted class for model training - it is notoriously difficult to gather enough examples for the model training, so when you get true negatives, celebrate!
The F1 score (also called the harmonic mean) is a single metric that combines the goals of precision and recall together and conveys the balance between both metrics. The equation is 2*((precision*recall)/(precision+recall)).
Let’s go back to our apple orange classifier. Imagine we have two oranges and two apples, and we have trained a model to classify fruits as either orange or apples. If the model predicts each of the four correctly, the model gets a 100 on precision, 100 on recall, and therefore an F1 score of 100. If the model instead predicts the two oranges correctly, predicts one apple as an orange, and doesn’t provide a label for the final apple (2 true positives, 1 false positive, and 1 false negative), the precision would be 75% (2 TP / (2 TP + 1 FP)), the recall would be 50% (2 TP / (2 TP + 1 FP + 1 FN), and the F1 score would be 60%. Here’s the calculation to follow along: 2 * ((0.75 * 0.5) / (0.75 + 0.5)) = 2 * (0.375 / 1.25) = 2 * 0.3 = 0.6 => 60%
These scores can be calculated on a class by class level, or across all of the classes. As you can see below, you can use the dropdown to look at the metric for a given class or select All to view the metric across all of your classes, which is an average for each class's metrics.
Confusion Matrix
A clean and unambiguous way to present prediction results for a classification model is to use a confusion matrix, also known as a contingency table.
For binary classification problems (models with two classes, e.g. “Cancer” and “Not_Cancer”) will have a table with 2 rows and 2 columns. Across the top are the observed classes, and down the side are the predicted class labels. Each cell contains the number of predictions made by the classification model that fall into that cell.
The aim is to have the majority of your predictions land in the “true positive” and “true negative” piles, so having the highest numbers found in the cells along the diagonal from the top left cell to the bottom right cell (see the image below).
In the example above, we have a multi-labeled classifier, with three classes (review, clinical_trial, and editorial). You can see that the majority of the predictions were correct, but of those that were wrong, we can see where the model got confused (and likely where we want to focus attention for the next model iteration). Here, we can see many “review” articles were misconstrued as “clinical_trials” (n = 253), whereas the model had very little trouble seeing an editorial and mislabeling it as a clinical trial (n = 16). Sometimes, though, it would see a clinical_trial and think it’s an editorial (n = 79), so the clinical trial class may need more examples to present in training to help the classifier nail down that class.
Average Loss
The average loss helps to evaluate and diagnose how well the model is learning. This includes all of the considerations of the optimization process, such as overfitting, underfitting, and convergence.
Loss is a value that represents the summation of errors in our model. It measures how well (or bad) our model is doing. If the errors are high, the loss will be high, which means that the model is not doing a good job. Otherwise, the lower it is, the better our model works.
To calculate the loss, a loss or cost function is used. There are several different loss functions to use. Seeing the loss over time can yield interesting findings for our models. If the loss value is not decreasing, but it just oscillates, the model might not be learning at all. However, if it’s decreasing in the training set but not in the validation set (or it decreases but there’s a notable difference), then the model might be overfitting. In other words, it might be overlearning from the training examples, becoming useless when presented with new examples. If that’s the case, you may wish to consider additional steps or a different approach: regularization, simpler models, or, even just reducing the learning rate.
How to Create a New Classification Model
1. In your “Models” view, click the “Actions” button
2. Select “Create Classification Model”
You will be taken to a page that will help guide you through how to define various parameters needed to build a classification model.
Step 1. Adding Classes & Training Data
1. Click “Add Class” in the “Step 1: Add Classes” panel
A new modal will pop up for adding your class. Provide the label you want for this class (e.g. “review” for systematic literature reviews)
2. Next, you’ll want to narrow your view to display only documents you want to be reviewed. This will build out a ground truth training dataset for the classification model to train on. User various filters as seen in Document Sets to assist you in this process. Check out the Boolean tips page to learn more about conducting searches for narrowing the document set.
Example
In the photo above, we used an annotation provided through PubMed that identifies a document as a review, systematic review, editorial, etc. We used the values “review” and “systematic review” to make a class with ~61K document samples (see bottom left for the total number of documents in your class’ set).
3. When your class set is ready, click “Save” to save those documents into the training dataset for the class (in this example, the class is “review”)
4. Continue to add however many classes as you need for your classification model. When ready, we’ll proceed to the next step: customizing parameters
💡 Pro Tip - It’s best to keep the number of examples for each class as close to equal as possible (this is known as balancing your classes). This ensures that your training run is not biased towards classes with a larger number of examples.
Step 2. Providing Customized Parameters
Once you have solidified your classes, you can adjust any model parameters used during the training. We also have defaults if you want to skip this step and use the base parameters instead.
See Classification Model Parameters for descriptions of each advanced parameter option.
Step 3. Commence Build!
You're all set to submit the classification model job. Push that button!
1. Click the “Train Model” button in the third panel
Upon clicking “Train Model”, you’ll be prompted to name your new model.
2. Click “Save”
Your model will begin automatically building, using the training datasets and parameters you have defined. You can leave or close out of this page and return at any time by clicking on your classification model’s name from the "Models" tab.
Note: If you plan to do multiple iterations of the model, we suggest having some versioning in your name (e.g. “Publication Classification V1”, “Publication Classification V2”, etc.).
When the model is done building, this page will automatically update to show you the model’s performance results, which were run on the test set you determined in your parameters portion of the training (see: Test Set Size in Classification Model Parameters). You will see metrics for F1 score, precision, and recall, as well as a confusion matrix and average loss function visualization.
To learn more about these performance metrics, please see Classification Model Performance Metrics.
How to Publish / Unpublish a Classification Model
Once you are satisfied with your model, you can publish it to be used for classification in Certara.AI Curate.
How to Publish a Classification Model
1. Click on the model you wish to publish.
2. Click the “Actions” button
3. select “Publish”
You’ll now notice that your model has gone from DRAFT to PUBLISHED status! You can now see it in the dropdown options available when you run Classify Documents, too.
How to Unpublish a Classification Model
1. To unpublish, you can click the Actions button and
2. Select the Unpublish option, which is only visible to published models
Your model will change from PUBLISHED to DRAFT and will no longer be available to use when running the Ask a Question modal. Any questions that were run using the model, however, will still have their predictions for that document set.
How to Retrain a Classification Model
If you would like to retrain a previously built model, either by adding new training data or by tinkering with custom parameters (or both!), here is how:
1. Click on the model you’d like to use as your base model (e.g. “Publication Type Classification Model v2”) You can review the following link to Create a Classification Model.
2. Click on the “Actions” button located in the top right corner
3. Click the “Retrain Model” option.
4. Adjust your dataset or custom parameters. See Classification Model Performance Metrics for descriptions of each advanced parameter option.
For example, since my classes are small, I found more documents for each class. I also increased my test set size from 10% to 30% and increased the number of epochs from 5 to 10.
5. Click “Train Model” and save a new name for this model (“Publication Type Classification v3”)
6. Once the model has been built, compare the results between this model’s performance metrics and the prior model’s metrics. Continue to tinker until you are satisfied with your model performance, at which point you can publish your model for use within the software (see How to Publish / Unpublish a Classification Model)
Tips for Natural Language Question Answering
Tip #1. Ask a formal question.
Good Example: "What is a biotech company in Boston?"
Bad Example: "Biotech companies in Boston" or "Show me biotech companies in
Boston"
Tip #2. Address the system as if you were querying a single document and ask for singular examples.
Questions are taken very literally.
Good Example: If you ask, "What is an analog for GLP-1-R?" you'll get "liraglutide"
Bad Example: If you ask the question "What are analogs for GLP-1- R?" you might get answers like "long-lasting".
Good Example: "What is a drug in clinical trials for the treatment of COVID-19?"
Bad Example: "What are drugs in clinical trials for the treatment of COVID-19?"
So, play around with wording if you find yourself getting some snarky answers back.
Tip #3. Avoid Qualifiers or a request for an opinion.
Good Example: "What is a company that is researching drugs for COVID-19?"
Bad Example: "What companies are most successful so far in finding effective.
drugs for COVID-19?"
Tip #4. Ask a single, simple question instead of a compound question.
Good Example: "What company produces Albuterol?"
Bad Example: "What is Albuterol and who produces it?"
Tip #5. Avoid time-based questions and use the date filter instead.
Good Example: "What is a company that has moved to Boston?" filter data for "search by date published".
Bad Example: "What companies have recently moved to Boston?"
Tip #6. Be as specific as possible.
Good Example: "What is a company that is researching drugs for the treatment of COVID-19?"
Bad Example: "What companies are working on Covid-19?"
Tip #7. To force the system to require a certain word or phrase to be present in search results, use single or double quotes around the word or phrase.
Example: "What is a venture capital company in Boston?"
Forced Behavior: “What is a ‘venture capital’ company in Boston?”





























































































