Wire
Nucleosome dataset opens 1.52M DNA examples
A new nucleosome-condensability dataset exposes 1,521,325 labeled DNA sequences, comprising 1,369,192 training examples and 152,133 validation examples in its live Hugging Face repository. Each record pairs a 147-base sequence with two condensability scores and a 147-position occupancy vector, but the card declares no license. For scientific-AI teams building beyond single-variant genomics prediction, the release offers a million-scale base-resolution task—provided they secure usage rights before treating public availability as permission.