| 📋 5,347 | 🧠 40+ | 🏢 1,500+ | 🛡️ 0.6% | 📊 7 | 💰 99.4% |
|---|---|---|---|---|---|
| Job descriptions | Canonical tech skills | Hiring companies | Actually disclosed salaries | Premium modules | Transparent salary modeling |
- What skill patterns are associated with higher modeled compensation in this dataset?
- How much of the job market data online is actually factual versus blindly modeled?
- What is the precise tech-stack difference between Product-tier and Consulting-tier companies?
- Which technologies have the strongest co-occurrence frequency (e.g. AWS + Snowflake)?
- What is my exact learning roadmap to break the ₹16L compensation ceiling?
- Cloud Multiplier (Model-Based): The salary estimation model assigns 1.15×–1.25× premiums for Cloud tools (Snowflake, Databricks, dbt), sourced from AmbitionBox/Glassdoor benchmarks. Note: 0 of 33 disclosed salaries in this dataset included cloud skills, so this is a model assumption.
- The "Dirty Data" Truth: Only 33 out of 5,347 (0.6%) job descriptions possessed explicitly disclosed salaries, showing that salary disclosure is extremely sparse in this dataset.
- Product Tier Premia (Model-Based): The model applies a 1.25× multiplier for Product-tier companies vs 1.10× for Consulting-tier (based on external benchmarks). 0 of 33 disclosed salaries came from either tier, so this is an assumption, not an observed finding.
- Top Tech Target: General programming capability (Python) combined with heavy data manipulation (SQL) retains absolute market dominance across 31% of total listings.
Data Honesty: Every metric displayed traces strictly to the custom Python NLP pipeline. Only 0.6% of salaries are observed (disclosed); all remaining salary figures are model-estimated using external benchmarks and clearly labelled. Salary model parameters (skill premiums, tier multipliers) are transparent assumptions, not disguised as observed data.
| Layer | Technology |
|---|---|
| Dashboard Framework | Vite |
| Frontend Mechanics | Vanilla JavaScript (ES6) |
| Visualizations | Chart.js 4.x + Glassmorphism Design System |
| Data Processing | Python · Pandas · NumPy |
| Extraction Pipeline | Regex + Custom Entity-Matching Hash Tables |
| Data Transfer Payload | Pre-compiled multidimensional JSON blocks |
The entire system is uncoupled. The heavy Python processing runs offline to calculate
dashboard_data.json, allowing the Vite web application to render with zero latency.
| Page | What It Shows |
|---|---|
| 🎛️ Command Center | Live market telemetry, top employer tracking, and our audited Trust badge |
| 🎯 Skill Demand Radar | Visual prevalence mapping mapping DBs vs Analytics vs Programming toolkits |
| 💰 Salary Intelligence | Simulator projecting lifetime LPA trajectory across 5 experience tiers |
| 🏢 Company War Room | Target searching across 1,500+ active hiring entities instantly |
| 🔗 Skill Synergy Map | Co-occurrence patterns showing which software combinations appear together most frequently |
| 🗺️ Career Pathfinder | Checkbox assessment that calculates the single missing tool driving the most ROI |
| 📰 Market Pulse | Automated, export-ready executive briefings |
5,347 total scraped Analyst & Data Scientist JD records
41 canonical target tools monitored (SQL, Tableau, dbt, etc.)
~200+ semantic synonyms collapsed via rigorous text processing
33 jobs containing disclosed, factual compensation (0.6%)
5,314 jobs completed via algorithmic proxy benchmarks (99.4%)
5,116 rows with explicitly identifiable corporate entities
git clone https://github.com/Yashaswini-V21/TalentPulse-Engine.git
cd TalentPulse-Engine/src1. Launch the Dashboard (Vite):
cd webapp
npm install
npm run dev
# Dashboard launches at http://localhost:51732. Rebuild the Intelligence Core (Python):
# Return to /src
pip install -r requirements.txt
python build_pipeline.py
python enrich_salary.py
python build_dashboard_json.py # compiles output payload to /webappTalentPulse_Engine/
│
├── 📁 src/
│ ├── 📄 build_pipeline.py ← Primary NLP dataset parsing
│ ├── 📄 enrich_salary.py ← Honest proxy-filling script (safeguards real data)
│ ├── 📄 build_dashboard_json.py ← Package compiler pushing CSVs to JSON payload
│ ├── 📄 requirements.txt ← Pinned Python dependencies
│ │
│ ├── 📁 data/
│ │ ├── raw/ ← Origin Dataset CSVs
│ │ └── clean/ ← Output analysis aggregations
│ │
│ ├── 📁 nlp/
│ │ └── skill_extractor.py ← Vocabulary logic mapping and custom dictionaries
│ │
│ ├── 📁 sql/
│ │ └── practice_queries.sql ← Embedded analytical logic equivalents
│ │
│ └── 📁 webapp/ ← Ultra-light frontend application
│ ├── 📄 index.html ← Entrypoint
│ ├── 📄 main.js ← Handles 7 distinct visual modules
│ ├── 📄 style.css ← Native glassmorphism token variables
│ └── 📁 public/
│ └── dashboard_data.json ← Pre-compiled multidimensional JSON from Python
│
├── 📁 assets/ ← Premium dashboard screenshots
└── 📄 README.md ← Project overview
Yashaswini V · LinkedIn · GitHub





