daVinci-Dev (Agent-native Trajectories)
daVinci-Dev (Agent-native Trajectories)
2,175 results
daVinci-Dev (Agent-native Trajectories)
Dementor — matrix model responses Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study. Companion to: Code…
The Vuk'uzenzele South African Multilingual Corpus
Clone from "friedrichor/MSR-VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K-A split protocol,…
X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq…
Dataset Card for FishBase Snapshots of FishBase data tables used by the rOpenSci package rfishbase and the FishBase…
FlyRank Internship — Warehouse Star Schema (Pseudonymized, Gated)
A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005…
XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar)…
osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions…
Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from…
ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality…
This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia…
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper…
Core-S2L1C Contains a global coverage of Sentinel-2 (Level 1C) patches, each of size 1,068 x 1,068 pixels. Source…
This repo contains (up to) 30k samples of 21 languages (top 20 languages by StackOverflow survey, html/css was…
Romanian Real Estate Listings Dataset
Hate Speech and Offensive Language