docs/skills.md
llms.txt, llms-full.txt, and per-topic markdown files designed for agent consumption); usage goes through standard Hugging Face APIs. Includes a bundled fetch_catalog.py script for filtered access by topic, type (datasets/models/blogs/spaces), or free-text search, plus reference guides for loading datasets via the datasets library (with streaming for billion-token corpora), running models via transformers or the Inference API/Providers (handling trust_remote_code requirements for custom architectures like Evo-2), and calling Spaces via gradio_client (with a worked BoltzGen example for protein binder design). Authenticates gated resources via HF_TOKEN from .env. Use cases: discovering the right dataset/model for a scientific ML task without trawling the broader Hub, fine-tuning on curated scientific data, citing methodology blogs from dataset/model authors, running interactive scientific demos (binder design, theorem proving, weather modeling) without local GPU setup, and bridging from "I need a model for protein/genome/molecule/climate/materials/astronomy" to working codedx/dxpy APIs. Covers secure authentication and project context, resumable data transfers, safe metadata/archival/deletion operations, Ubuntu 24.04 app and applet development, modern region-scoped dxapp.json configuration, jobs and workflow analyses with cost/retry/resource controls, native workflows, WDL/CWL through dxCompiler, and licensed Nextflow imports. Includes offline validators for manifests and installed SDK compatibilityLABARCHIVES_* variables; remote writes require a reviewed, redacted plan and explicit approvalLPath data operations plus LatchFile/LatchDir, transactional Registry APIs and samplesheets, CPU/GPU resource sizing, UI metadata and launch plans, staging and programmatic execution, Nextflow and version-aware Snakemake packaging, ready-to-use workflow discovery, automations, and the OAuth-based Latch MCPnextflow.config, configuring executors and containers (Docker, Singularity/Apptainer, Conda, Wave), scaling to HPC/SLURM or cloud (AWS Batch, Google Batch, Azure, Kubernetes), and debugging failed or -resume runs. Use for any reproducible scientific/bioinformatics workflow work and for authoring nf-core-compliant pipelines, modules, configs, and linting--execute is explicit, and never execute mutations; create/update/publish operations remain non-executing plans for authorized reviewlimit for top-N links, and diy() for custom regressors. Core engine for pySCENIC (cisTarget pruning and regulon activity are downstream). Use for bulk or single-cell expression matrices when inferring TF–target regulatory links from co-expression patternsget.aggregate(), trajectory inference (PAGA, diffusion maps), and visualization. Key features include: efficient handling of large datasets using sparse matrices and experimental Dask out-of-core support, integration with scvi-tools for advanced analysis, batch correction methods (ComBat), and publication-quality plotting. Optional GPU acceleration via rapids-singlecell. Use cases: single-cell RNA-seq analysis, cell-type identification, exploratory cluster markers, pseudobulk DE workflows (with pydeseq2), trajectory analysis, and comprehensive single-cell genomics workflowszarr.codecs compression (Blosc, gzip, zstd), partial chunk reads, consolidated metadata, sharding, and integration with NumPy, Dask, and Xarray. Use for out-of-core arrays, cloud-native pipelines, and large scientific datasets (genomics, imaging, climate). Skill: zarr-pythonln.track()/ln.finish() and @ln.flow()/@ln.step(), schema validation for DataFrame/AnnData/SpatialData/TileDB-SOMA, and ontology-backed annotation with Bionty 2.x. Includes current guidance for projects, branches, spaces, storage backends (local, S3, GCS, S3-compatible, Hugging Face), SQLite/Postgres deployment, safe credential handling, external-data validation, and integrations with Nextflow/nf-lamin, Snakemake, Redun, W&B, MLflow, Lightning, scVI-tools, DuckDB, and Vitessce..map() for batch processing, input concurrency and dynamic batching for I/O-bound workloads, Sandboxes for isolated execution of untrusted or agent-generated code with network (CIDR) restrictions, and resource configuration (CPU cores, memory, ephemeral disk up to 3 TiB). Supports custom Docker images, Micromamba/Conda environments, integration with Hugging Face/Weights & Biases, and distributed multi-GPU training. Free tier includes $30/month credits. Use cases: ML model deployment and inference (LLMs, image generation, speech, embeddings), GPU-accelerated training and fine-tuning, batch processing large datasets in parallel, scheduled compute-intensive jobs, serverless API deployment with autoscaling, protein folding and computational biology, scientific computing requiring distributed compute or specialized hardware, and data pipeline automationdeepchem[torch], [tensorflow], [jax]). Use for ADMET/toxicity prediction, materials properties, and transfer learning on small datasetspkg_resources; dataset, checkpoint, benchmark, and remote-oracle network/storage effects require review and explicit approvalesm SDK workflows for local ESM3/ESMC open models, Forge-hosted ESM3 and ESMC inference with ESM_API_KEY, Biohub-hosted ESMC embeddings, and ESMFold2 all-atom structure prediction. Use cases: novel protein design, sequence/structure co-design, protein embeddings, function annotation, variant generation, and directed evolution workflowsx-api-key header) or MCP server — no local GPUs required. Covers structure prediction (AlphaFold, Boltz-2, Chai-1, ESMFold2), protein/binder/de novo design (RFdiffusion, ProteinMPNN/LigandMPNN, BoltzGen, BindCraft), antibody and nanobody design, humanization and developability, protein-ligand docking (DiffDock, AutoDock Vina) and binding-affinity prediction, MSA generation, and molecular dynamics — all through one uniform job API with single and batch submission and tool chaining (design → fold → score). Discovers tools and schemas live (GET /tools, MCP getAvailableTools/getJobSchema) and reads the key from TAMARIND_API_KEY. There is no official Python SDK — use plain requests or the MCP server. Use cases: cloud structure prediction and protein design without provisioning hardware, high-throughput batch characterization of sequences/designs, and chaining design-fold-score pipelinespufferlib==3.0.0 provides the Python/Gymnasium/PettingZoo emulation, vector, and Torch PuffeRL surface; the redesigned native 4.0 source line is incompatible and removes those 3.0 modules. Start with CPU-only synthetic tools and audit native builds, plug-ins, assets, checkpoints, GPU use, logging, and network effects before executionshap.Explanation API, output-aware Tree/Linear/Permutation/Partition/Deep explainers, background-data and masker selection, multiclass slicing, additivity validation, cohort analysis, text/image workflows, and current bar/beeswarm/waterfall/scatter visualizations. Emphasizes reproducible reports and clear limits: predictive attributions are not causal effects, fairness certificates, or recourseData/HeteroData, 60+ conv layers (GCN, GAT, GraphSAGE, GIN), node/link/graph classification, heterogeneous graphs, neighbor sampling (NeighborLoader, LinkNeighborLoader), OGB and built-in datasets, custom dataset loading, GNN explainability, and scaling via DDP, Lightning, and torch.compilepymatgen==2026.5.4, pymatgen-core==2026.7.16, and mp-api==0.46.4 on Python 3.11+. Covers provenance-preserving local phase diagrams, symmetry sensitivity, transformations, and electronic-structure I/O; Materials Project queries are explicitly bounded and require user-approved network access plus the named MP_API_KEYTreePattern matching, Robinson-Foulds comparison and distance matrices, PhyloTree duplication/speciation inference and gene-tree reconciliation, local NCBI and GTDB taxonomy databases, browser-based SmartView exploration, and PNG/PDF/SVG rendering. Includes ETE 3→4 migration guidance and validated command-line helpers. Use dedicated alignment and tree-inference tools before ETE when starting from raw sequencestree.py state manager (init/observe/add-node/set-evidence/propagate/prune/merge/status/validate) that keeps the tree consistent and auditable, plus reference docs on the HTR methodology, the executor brief, the final-report template, and how to run the upstream arbor CLI instead. Use cases: "improve my model's eval score in fewer steps", optimizing an agent or search harness for higher pass rate/accuracy, tuning a data-generation pipeline judged by downstream model behavior, MLE-bench / Kaggle-style "improve the submission" tasks, and any long-horizon "make this artifact better and don't just memorize the dev set" workflow that needs structured, branching exploration with an audit traileverything(), full-text content indexing, saved search management, and file upload/download. Optional CLI and built-in MCP server (pyzotero 1.12+) for searching local Zotero 7 libraries including full-text PDF search and Semantic Scholar integration. Use cases: building research automation pipelines that integrate with Zotero, bulk importing references, exporting bibliographies programmatically, managing large reference collections, syncing library metadata, enriching bibliographic data, and connecting LLM agents to a local Zotero library..pptx from strict author-approved local content and asset manifests using python-pptx 1.0.2. It does not use HTML conversion, network services, external templates, or generated claims. Exact physical size, printer rules, provenance, asset hashes/licenses, accessibility, package security, and approval hashes must pass before generation; final PowerPoint/PDF/printer/author review remains manualword/document.xml → zip, run merging so text is findable in the XML, tracked-change redlining with author-scoped validation, accept_changes.py, and a comment helper that writes all six cross-linked comment parts. Includes OOXML schema validation and a LibreOffice wrapper for render-and-inspect verification. Use for reports, memos, letters, templates, and any Word document deliverable. Created and maintained by Anthropicresult.markdown API; safe local/batch/literature workflows; the official vision OCR plugin; Azure Document Intelligence and Content Understanding; custom plugins; and the local MarkItDown MCP server. Includes explicit SSRF, archive, prompt-injection, plugin, credential, external-transcription, and cloud-data controls, plus deterministic local-input helper scripts with an external-service gatetext_items (coordinates, font metadata, OCR confidence). Built-in Tesseract OCR with optional HTTP OCR servers (EasyOCR/PaddleOCR-compatible API), page subsets, encrypted PDFs, stdin/bytes parsing, and PNG page screenshots for multimodal agents. CLI (lit parse, lit batch-parse, lit screenshot) and Python API (LiteParse, search_items). Targets liteparse 2.0.0, Python 3.10+. Use when you need spatial grounding for RAG, batch literature ingestion, or agent vision—not for Markdown (MarkItDown) or PDF merge/split/forms (pdf skill)add_slide.py), orphan cleanup (clean.py), thumbnail grids for layout selection, and OOXML validation that catches the chart and slide defects PowerPoint refuses to open. Includes slide design guidance (palettes, typography, spacing) and a required content/file/visual QA pass. Use for slide decks, pitch decks, extracting text from presentations, editing existing slides, and working with templates, layouts, speaker notes, and comments. Created and maintained by Anthropicrecalc.py (openpyxl writes formulas with no cached values), which formulas survive that recalculation (_xlfn. prefixes; never XLOOKUP/FILTER/UNIQUE), openpyxl gotchas (two-pass reading, destructive data_only=True saves, merged cells, keep_vba), external-link loss on re-save, and financial-model conventions for color coding, number formats, and assumption structure. Created and maintained by Anthropichypogenic==0.3.5/HypoRefine workflows over labeled text datasets, task configs, hypothesis banks, and HypoBench data. The software proposes candidate textual patterns and task-prediction statistics; held-out accuracy is not experimental confirmation, causal evidence, novelty, or scientific validity. Model providers, Redis, credentials, data, and network use require separate approval, and model calls never start automaticallycategory="research paper" plus academic domain allowlists, and batch URL extraction