SynthDet - An end-to-end object detection pipeline using synthetic data
-
Updated
Dec 5, 2024 - C#
SynthDet - An end-to-end object detection pipeline using synthetic data
Generate conversational, tool-calling, structured-output, and preference datasets — easily and at scale
Deterministic synthetic data generator for realistic, correlated, and noisy test records across 68 locales. Rust CLI/Python/Node.js/Browser WASM/Go/PHP/Ruby/MCP
The MERIT Dataset is a fully synthetic, labeled dataset created for training and benchmarking LLMs on Visually Rich Document Understanding tasks. It is also designed to help detect biases and improve interpretability in LLMs, where we are actively working. This repository is actively maintained, and new features are continuously being added.
Efficient and multi-language generation from context free or sensitive grammars (CFG/CSG)
PhishNet is an experimental research project implementing Reinforced Self-Training (ReST) human-aligned with crafted instructions and fine-tuned models to craft a high-quality synthetic dataset of phishing emails.
50 URS + 50 paired FS for GxP-regulated computer systems. CC-BY-SA 4.0. Synthetic, AI-authored, not a regulated record.
Synthetic data generator for hallucinating entire companies as browsable file shares of realistic business docs & emails with a deterministic ground-truth answer key. Generates: .docx/.pdf/.xlsx/.pptx/.eml, and pre-2007 binaries. Used for augmenting research data sets and testing agentic systems against meaningful synthetic examples.
Synthetic sounds datasets and real sounds datasets of waterflow sounds for the repo 'Neural-Texture-Sound-Synthesis-with-physically-driven-continuous-controls'.
This repository contains a synthetic dataset and a step-by-step exploratory data analysis (EDA) workflow for a classification problem simulating customer churn prediction. The dataset is fully generated to mimic real-world scenarios with numerical, categorical, and binary target variables.
An open-source software for synthetic web-based user interface and content dataset generation. To cite this Original Software Publication: https://www.sciencedirect.com/science/article/pii/S2352711022000073
Synthetic Dataset Generation - GANS
SaaS Email Deliverability & Abuse Intelligence — star-schema analytics project: a mid-year deliverability incident and a Free-tier abuse cluster, with Python data pipeline, DuckDB SQL, and a Power BI dashboard.
IteraBeast is a synthetic data generation engine designed to simplify and accelerate the process of generating large-scale synthetic datasets for training and fine-tuning LLMs.
Synthetic satellite stereo dataset with high-precision ground-truth disparity, rendered in Unreal Engine 5. 10,000 disparity sets across 8 photo realistic virtual 3D environments. CVPR Workshop EarthVision 2026.
Synthetic underwater acoustic waveforms for self-supervised learning. 12,000 5-second clips at 16 kHz covering 4 vessel classes + no-vessel ambient. Non-overlapping shaft rates, blade-gated cavitation bursts, Knudsen-model sea noise. License: TBD by Altair Infrasec Pvt. Ltd. Contact styagi@oravontsystems.com for details
A spec-driven Databricks pipeline for a fictional Austin lawn-care company - one simulated quarter of OpenSpec practice. Synthetic history, generated for a write-up.
Free synthetic industrial-manual RAG dataset, 14-test benchmark, TCO workbook, and buyer tools for bounded private AI pilots
Synthetic training-data generation and model training for the comic-localizer text-mask segmenter
Normalizing-flow generation of realistic synthetic clinical outcomes
To associate your repository with the synthetic-dataset topic, visit your repo's landing page and select "manage topics."