Translate project documentation to English
This commit is contained in:
@@ -1,10 +1,10 @@
|
|||||||
# SciMesh
|
# SciMesh
|
||||||
|
|
||||||
Минимальный локальный поиск молекул ChEMBL, похожих на гефитиниб (`CHEMBL939`).
|
Minimal local search for ChEMBL molecules similar to gefitinib (`CHEMBL939`).
|
||||||
|
|
||||||
Скрипт находит `CHEMBL939` в TSV и использует его `canonical_smiles` как эталон. Затем во втором потоковом проходе по файлу строит Morgan fingerprints (`radius=2`, `fpSize=2048`) и ранжирует остальные валидные SMILES по Tanimoto similarity. Невалидные SMILES и сам `CHEMBL939` пропускаются. В памяти остаются только 20 лучших результатов (или значение `--top`).
|
The script finds `CHEMBL939` in the TSV file and uses its `canonical_smiles` as the reference. It then makes a second streaming pass through the file, generates Morgan fingerprints (`radius=2`, `fpSize=2048`), and ranks the remaining valid SMILES by Tanimoto similarity. Invalid SMILES and `CHEMBL939` itself are skipped. Only the best 20 results (or the value passed to `--top`) are kept in memory.
|
||||||
|
|
||||||
## Установка
|
## Installation
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python -m venv .venv
|
python -m venv .venv
|
||||||
@@ -12,45 +12,45 @@ source .venv/bin/activate
|
|||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
```
|
```
|
||||||
|
|
||||||
RDKit также можно установить через conda-forge:
|
RDKit can also be installed through conda-forge:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
conda install -c conda-forge rdkit
|
conda install -c conda-forge rdkit
|
||||||
```
|
```
|
||||||
|
|
||||||
## Запуск
|
## Usage
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scimesh.py chembl_37_chemreps.txt -o gefitinib_similarities.csv
|
python scimesh.py chembl_37_chemreps.txt -o gefitinib_similarities.csv
|
||||||
```
|
```
|
||||||
|
|
||||||
По умолчанию создаётся CSV с колонками `rank,chembl_id,canonical_smiles,similarity` и выводится тот же top-20 в терминал. Для другого размера выборки:
|
By default, the script writes a CSV with `rank,chembl_id,canonical_smiles,similarity` columns and prints the same top 20 results to the terminal. To choose a different number of results:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scimesh.py chembl_37_chemreps.txt --top 50 -o top_50.csv
|
python scimesh.py chembl_37_chemreps.txt --top 50 -o top_50.csv
|
||||||
```
|
```
|
||||||
|
|
||||||
Во время поиска статус выводится в `stderr` каждые 100 000 строк: количество обработанных строк, текущая и средняя скорость, прошедшее время и число пропущенных невалидных SMILES. Период можно изменить или отключить:
|
During the search, status is written to `stderr` every 100,000 rows: number of processed rows, current and average rates, elapsed time, and the number of invalid SMILES skipped. The interval can be changed or disabled:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scimesh.py chembl_37_chemreps.txt --progress-every 500000
|
python scimesh.py chembl_37_chemreps.txt --progress-every 500000
|
||||||
python scimesh.py chembl_37_chemreps.txt --progress-every 0
|
python scimesh.py chembl_37_chemreps.txt --progress-every 0
|
||||||
```
|
```
|
||||||
|
|
||||||
## Быстрый тест на части базы
|
## Quick test on part of the database
|
||||||
|
|
||||||
Опция `--max-rows` ограничивает второй проход первыми `N` строками TSV. Сам `CHEMBL939` перед этим всё равно находится в отдельном потоковом проходе, поэтому эталон остаётся тем же. Получившийся CSV — это top‑20 только по обработанной части, а не по полной базе.
|
The `--max-rows` option limits the second pass to the first `N` TSV rows. `CHEMBL939` is still found in its own streaming pass first, so the reference stays the same. The resulting CSV is the top 20 only within the processed subset, not the full database.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scimesh.py chembl_37_chemreps.txt --max-rows 10000 -o test_results.csv
|
python scimesh.py chembl_37_chemreps.txt --max-rows 10000 -o test_results.csv
|
||||||
```
|
```
|
||||||
|
|
||||||
## Изображения структур
|
## Structure images
|
||||||
|
|
||||||
Передайте каталог в `--images-dir`, чтобы создать `CHEMBL939_gefitinib.png` с гефитинибом и `top_candidates.png` с сеткой top‑кандидатов. На изображениях кандидатов указаны ранг, ChEMBL ID и Tanimoto similarity.
|
Pass a directory to `--images-dir` to create `CHEMBL939_gefitinib.png` for gefitinib and `top_candidates.png` with a grid of top candidates. Candidate images show rank, ChEMBL ID, and Tanimoto similarity.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scimesh.py chembl_37_chemreps.txt --images-dir structures
|
python scimesh.py chembl_37_chemreps.txt --images-dir structures
|
||||||
```
|
```
|
||||||
|
|
||||||
Опция `--image-columns` задаёт число структур в строке сетки (по умолчанию `4`).
|
The `--image-columns` option controls the number of structures per grid row (default: `4`).
|
||||||
|
|||||||
+1
-1
@@ -1,5 +1,5 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""SciMesh: потоковый поиск молекул ChEMBL, похожих на гефитиниб (CHEMBL939)."""
|
"""SciMesh: streaming search for ChEMBL molecules similar to gefitinib."""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user