biotech Deep-Interact Studio

User guide

Manual

A compact guide to using Deep-Interact Studio, from preparing a CSV to comparing trained models.

  1. Choose a task
  2. Load data
  3. Map columns
  4. Configure the model
  5. Train
  6. Check results
  7. Run inference

What this webtool does

Deep-Interact Studio builds sequence-based interaction classifiers for biological pair prediction. It turns molecules and sequences into embeddings, trains a classifier, reports validation metrics, and lets you reuse completed models for inference. The app supports three tasks:

hub

PPI

Protein-Protein Interaction

Predict whether two protein sequences interact.

medication

DTPI

Drug-Target Protein Interaction

Predict whether a compound binds a protein target.

genetics

RPI

RNA-Protein Interaction

Predict whether an RNA sequence binds a protein.

1 Prepare data

Use a CSV with two input columns and one binary label column. The label is 1 for interaction and 0 for no interaction. Column names do not need to match exactly because each model-building page lets you map your columns after upload.

TaskUse forExpected data
PPIPredict whether two protein sequences interact. ProteinA, ProteinB, lable
DTPIPredict whether a compound binds a protein target. SMILES, Protein, lable
RPIPredict whether an RNA sequence binds a protein. RNA, Protein, lable

Practical checks before training

  • Keep labels as 0 and 1.
  • Remove duplicate pairs where possible.
  • Use balanced positives and negatives for the first experiment.
  • Protein inputs are validated against the current page limit of 512 residues.
  • Start with a small sample if the CSV is large.

2 Train a model

Open the page for your task and follow the same basic sequence.

  1. Upload a CSV or load the example data.
  2. Map the sequence, molecule, and label columns.
  3. Choose how many pairs to use and the train/test split.
  4. Select the embedding model where the page offers options.
  5. Build a classifier from the available layers.
  6. Set training parameters such as epochs, batch size, and learning rate.
  7. Submit the job and save the Run ID and access token.
Recommended starting settings
SettingGood first choice
Dataset sample100-500 pairs for a first test
Train/test split80/20
Class balance50% positive, 50% negative
Protein embeddingESM2 35M when available
ClassifierTwo Linear layers with dropout
Epochs30-50 with early stopping
Learning rate0.001

3 Track results

After submission, use the Run ID to follow the job and inspect model quality.

  • Job Status lists submitted training and inference jobs.
  • Check Model Results shows training curves, metrics, dataset summary, downloads, and failures.
  • Main metrics to compare are AUROC, Average Precision, MCC, F1, and the confusion matrix.
How to read the main metrics
MetricWhat it tells you
AUROCHow well positives rank above negatives across thresholds
Average PrecisionBetter than AUROC when positives are rare
MCCBalanced single-score metric for imbalanced data
F1Balance between precision and recall at one threshold
AccuracyEasy to understand, but misleading when classes are imbalanced

4 Run inference

Use a completed training Run ID to predict new pairs. You can enter a single pair manually or upload a batch CSV. If labels are included in a batch file, the app also reports inference metrics.

5 Compare runs

Use comparison pages when you train multiple models or run multiple inference batches.

  • Compare Models compares completed training runs from the same task type.
  • Compare Inferences compares completed inference runs from the same task type.

FAQ & Troubleshooting

What are the job submission limits?

Training jobs are accepted only when they stay within these limits:

LimitCurrent value
Upload size per file100 MB
Total upload request size100 MB
Selected training pairsUp to 100,000 pairs
Model sizeUp to 5,000,000 trainable parameters
Protein sequence length on builder pages512 residues
Training wall-clock time4 hours, then the job is stopped
Training submissions10 training jobs per IP per 3 hours
Total active training queue20 queued or running training jobs platform-wide

If your job is rejected, reduce the selected positive/negative pair counts, simplify the architecture, trim long protein sequences, or wait for queued/running jobs to finish.

What does the training queue limit mean?

The queue limit is platform-wide. It counts all training jobs with status queued or running, not only your own jobs. When the active training queue reaches 20 jobs, new training submissions are temporarily blocked until at least one job completes, fails, or is cleaned up.

This protects a free shared research deployment from building a very long backlog. Your per-IP quota still applies separately: one IP can submit up to 10 training jobs in a rolling 3-hour window, provided the platform queue is not already full.

What are the inference limits?

Inference jobs have separate limits from training:

LimitCurrent value
Single-pair input fields512 characters each
Batch inference CSVUp to 60,000 pairs
Inference submissions20 inference requests per minute
Batch inference quota15 batch jobs per IP per 5 hours
Single-pair inference quota30 single-pair jobs per IP per 5 hours

For large inference files, keep only the required columns and remove duplicate rows before upload.

How many pairs should I use?

For a first run, use 100-500 balanced pairs to check that the workflow, column mapping, and model settings are correct. For stronger models, increase the pair count while watching runtime and class balance.

Very small datasets can overfit. Very large datasets can time out or exceed memory limits, especially with larger embedding models or large classifier architectures.

What should I save after submitting a job?

Save the Run ID and the access token immediately. The access token is needed to view results, download models, run inference with the model, and compare runs; it is shown once and cannot be recovered. Save the cancel token too if you may need to stop the job.

This browser remembers tokens for runs you submitted here, so you only need to paste a token when opening a run from a different browser, device, or web address.

A job failed

Check Job Status or Check Model Results for the error. Common causes are invalid sequences, invalid SMILES strings, too many selected pairs, GPU memory limits, or timeout. Reduce the sample size and retry with a smaller embedding model if the job is too heavy.

Which metrics should I report?

Report AUROC, Average Precision, MCC, and the confusion matrix. Accuracy can be included, but it should not be the primary metric when classes are imbalanced.

For candidate screening, also inspect precision and recall at the threshold you plan to use. Lower thresholds find more possible interactors; higher thresholds return fewer but more confident candidates.

Results look too good

Check for duplicate pairs and shared entities between train and test data. High overlap can make validation metrics look better than real-world performance. For publication-style evaluation, prefer disjoint splits where possible.

Accuracy is high but AUROC or MCC is poor

This usually means the dataset is imbalanced. Use AUROC, Average Precision, MCC, and the confusion matrix instead of relying only on accuracy.

How many runs can I compare?

The comparison pages accept up to 5 runs at a time. Compare only runs from the same task type so the metrics and prediction outputs are meaningful.