docs/skills.md
llms.txt, llms-full.txt, and per-topic markdown files designed for agent consumption); usage goes through standard Hugging Face APIs. Includes a bundled fetch_catalog.py script for filtered access by topic, type (datasets/models/blogs/spaces), or free-text search, plus reference guides for loading datasets via the datasets library (with streaming for billion-token corpora), running models via transformers or the Inference API/Providers (handling trust_remote_code requirements for custom architectures like Evo-2), and calling Spaces via gradio_client (with a worked BoltzGen example for protein binder design). Authenticates gated resources via HF_TOKEN from .env. Use cases: discovering the right dataset/model for a scientific ML task without trawling the broader Hub, fine-tuning on curated scientific data, citing methodology blogs from dataset/model authors, running interactive scientific demos (binder design, theorem proving, weather modeling) without local GPU setup, and bridging from "I need a model for protein/genome/molecule/climate/materials/astronomy" to working codedx/dxpy APIs. Covers secure authentication and project context, resumable data transfers, safe metadata/archival/deletion operations, Ubuntu 24.04 app and applet development, modern region-scoped dxapp.json configuration, jobs and workflow analyses with cost/retry/resource controls, native workflows, WDL/CWL through dxCompiler, and licensed Nextflow imports. Includes offline validators for manifests and installed SDK compatibilityLABARCHIVES_* variables; remote writes require a reviewed, redacted plan and explicit approvalLPath data operations plus LatchFile/LatchDir, transactional Registry APIs and samplesheets, CPU/GPU resource sizing, UI metadata and launch plans, staging and programmatic execution, Nextflow and version-aware Snakemake packaging, ready-to-use workflow discovery, automations, and the OAuth-based Latch MCPnextflow.config, configuring executors and containers (Docker, Singularity/Apptainer, Conda, Wave), scaling to HPC/SLURM or cloud (AWS Batch, Google Batch, Azure, Kubernetes), and debugging failed or -resume runs. Use for any reproducible scientific/bioinformatics workflow work and for authoring nf-core-compliant pipelines, modules, configs, and linting--execute is explicit, and never execute mutations; create/update/publish operations remain non-executing plans for authorized reviewlimit for top-N links, and diy() for custom regressors. Core engine for pySCENIC (cisTarget pruning and regulon activity are downstream). Use for bulk or single-cell expression matrices when inferring TF–target regulatory links from co-expression patternsget.aggregate(), trajectory inference (PAGA, diffusion maps), and visualization. Key features include: efficient handling of large datasets using sparse matrices and experimental Dask out-of-core support, integration with scvi-tools for advanced analysis, batch correction methods (ComBat), and publication-quality plotting. Optional GPU acceleration via rapids-singlecell. Use cases: single-cell RNA-seq analysis, cell-type identification, exploratory cluster markers, pseudobulk DE workflows (with pydeseq2), trajectory analysis, and comprehensive single-cell genomics workflowszarr.codecs compression (Blosc, gzip, zstd), partial chunk reads, consolidated metadata, sharding, and integration with NumPy, Dask, and Xarray. Use for out-of-core arrays, cloud-native pipelines, and large scientific datasets (genomics, imaging, climate). Skill: zarr-pythonln.track()/ln.finish() and @ln.flow()/@ln.step(), schema validation for DataFrame/AnnData/SpatialData/TileDB-SOMA, and ontology-backed annotation with Bionty 2.x. Includes current guidance for projects, branches, spaces, storage backends (local, S3, GCS, S3-compatible, Hugging Face), SQLite/Postgres deployment, safe credential handling, external-data validation, and integrations with Nextflow/nf-lamin, Snakemake, Redun, W&B, MLflow, Lightning, scVI-tools, DuckDB, and Vitessce..map() for batch processing, input concurrency and dynamic batching for I/O-bound workloads, Sandboxes for isolated execution of untrusted or agent-generated code with network (CIDR) restrictions, and resource configuration (CPU cores, memory, ephemeral disk up to 3 TiB). Supports custom Docker images, Micromamba/Conda environments, integration with Hugging Face/Weights & Biases, and distributed multi-GPU training. Free tier includes $30/month credits. Use cases: ML model deployment and inference (LLMs, image generation, speech, embeddings), GPU-accelerated training and fine-tuning, batch processing large datasets in parallel, scheduled compute-intensive jobs, serverless API deployment with autoscaling, protein folding and computational biology, scientific computing requiring distributed compute or specialized hardware, and data pipeline automationdeepchem[torch], [tensorflow], [jax]). Use for ADMET/toxicity prediction, materials properties, and transfer learning on small datasetspkg_resources; dataset, checkpoint, benchmark, and remote-oracle network/storage effects require review and explicit approvalesm SDK workflows for local ESM3/ESMC open models, Forge-hosted ESM3 and ESMC inference with ESM_API_KEY, Biohub-hosted ESMC embeddings, and ESMFold2 all-atom structure prediction. Use cases: novel protein design, sequence/structure co-design, protein embeddings, function annotation, variant generation, and directed evolution workflowsx-api-key header) or MCP server — no local GPUs required. Covers structure prediction (AlphaFold, Boltz-2, Chai-1, ESMFold2), protein/binder/de novo design (RFdiffusion, ProteinMPNN/LigandMPNN, BoltzGen, BindCraft), antibody and nanobody design, humanization and developability, protein-ligand docking (DiffDock, AutoDock Vina) and binding-affinity prediction, MSA generation, and molecular dynamics — all through one uniform job API with single and batch submission and tool chaining (design → fold → score). Discovers tools and schemas live (GET /tools, MCP getAvailableTools/getJobSchema) and reads the key from TAMARIND_API_KEY. There is no official Python SDK — use plain requests or the MCP server. Use cases: cloud structure prediction and protein design without provisioning hardware, high-throughput batch characterization of sequences/designs, and chaining design-fold-score pipelinespufferlib==3.0.0 provides the Python/Gymnasium/PettingZoo emulation, vector, and Torch PuffeRL surface; the redesigned native 4.0 source line is incompatible and removes those 3.0 modules. Start with CPU-only synthetic tools and audit native builds, plug-ins, assets, checkpoints, GPU use, logging, and network effects before executionshap.Explanation API, output-aware Tree/Linear/Permutation/Partition/Deep explainers, background-data and masker selection, multiclass slicing, additivity validation, cohort analysis, text/image workflows, and current bar/beeswarm/waterfall/scatter visualizations. Emphasizes reproducible reports and clear limits: predictive attributions are not causal effects, fairness certificates, or recourseData/HeteroData, 60+ conv layers (GCN, GAT, GraphSAGE, GIN), node/link/graph classification, heterogeneous graphs, neighbor sampling (NeighborLoader, LinkNeighborLoader), OGB and built-in datasets, custom dataset loading, GNN explainability, and scaling via DDP, Lightning, and torch.compilepymatgen==2026.5.4, pymatgen-core==2026.7.16, and mp-api==0.46.4 on Python 3.11+. Covers provenance-preserving local phase diagrams, symmetry sensitivity, transformations, and electronic-structure I/O; Materials Project queries are explicitly bounded and require user-approved network access plus the named MP_API_KEYTreePattern matching, Robinson-Foulds comparison and distance matrices, PhyloTree duplication/speciation inference and gene-tree reconciliation, local NCBI and GTDB taxonomy databases, browser-based SmartView exploration, and PNG/PDF/SVG rendering. Includes ETE 3→4 migration guidance and validated command-line helpers. Use dedicated alignment and tree-inference tools before ETE when starting from raw sequencestree.py state manager (init/observe/add-node/set-evidence/propagate/prune/merge/status/validate) that keeps the tree consistent and auditable, plus reference docs on the HTR methodology, the executor brief, the final-report template, and how to run the upstream arbor CLI instead. Use cases: "improve my model's eval score in fewer steps", optimizing an agent or search harness for higher pass rate/accuracy, tuning a data-generation pipeline judged by downstream model behavior, MLE-bench / Kaggle-style "improve the submission" tasks, and any long-horizon "make this artifact better and don't just memorize the dev set" workflow that needs structured, branching exploration with an audit traileverything(), full-text content indexing, saved search management, and file upload/download. Optional CLI and built-in MCP server (pyzotero 1.12+) for searching local Zotero 7 libraries including full-text PDF search and Semantic Scholar integration. Use cases: building research automation pipelines that integrate with Zotero, bulk importing references, exporting bibliographies programmatically, managing large reference collections, syncing library metadata, enriching bibliographic data, and connecting LLM agents to a local Zotero library..pptx from strict author-approved local content and asset manifests using python-pptx 1.0.2. It does not use HTML conversion, network services, external templates, or generated claims. Exact physical size, printer rules, provenance, asset hashes/licenses, accessibility, package security, and approval hashes must pass before generation; final PowerPoint/PDF/printer/author review remains manualresult.markdown API; safe local/batch/literature workflows; the official vision OCR plugin; Azure Document Intelligence and Content Understanding; custom plugins; and the local MarkItDown MCP server. Includes explicit SSRF, archive, prompt-injection, plugin, credential, external-transcription, and cloud-data controls, plus deterministic local-input helper scripts with an external-service gatetext_items (coordinates, font metadata, OCR confidence). Built-in Tesseract OCR with optional HTTP OCR servers (EasyOCR/PaddleOCR-compatible API), page subsets, encrypted PDFs, stdin/bytes parsing, and PNG page screenshots for multimodal agents. CLI (lit parse, lit batch-parse, lit screenshot) and Python API (LiteParse, search_items). Targets liteparse 2.0.0, Python 3.10+. Use when you need spatial grounding for RAG, batch literature ingestion, or agent vision—not for Markdown (MarkItDown) or PDF merge/split/forms (pdf skill)hypogenic==0.3.5/HypoRefine workflows over labeled text datasets, task configs, hypothesis banks, and HypoBench data. The software proposes candidate textual patterns and task-prediction statistics; held-out accuracy is not experimental confirmation, causal evidence, novelty, or scientific validity. Model providers, Redis, credentials, data, and network use require separate approval, and model calls never start automaticallycategory="research paper" plus academic domain allowlists, and batch URL extraction