Finalization, Refactoring and Validation of an Automation, Ranking and Data Analysis System with Machine Learning
EUR 250–750
About the project
Finalization, Refactoring and Validation of an Automation, Ranking and Data Analysis System with Machine Learning Project Summary This project consists of a partially developed system with a functional architecture that requires finalization, refactoring, validation, and deployment to operate in a stable, reliable manner with correct metrics. The system collects historical data, processes information, generates rankings, exports files by ID range, and provides an interface for visualization and analysis. The current code is functional but concentrated in a single monolithic file of approximately 15,600 lines, making maintenance and evolution unfeasible. What Already Exists (Confirmed Codebase) Backend: FastAPI (Python) with approximately 180 endpoints. Frontend: React + TypeScript + Vite, with functional dashboard. Database: MongoDB with modeled collections and validated historical data (over 7,000 records). Machine Learning: Pipeline with implemented models (Gradient Boosting, Random Forest, LSTM). Automation: Selenium for data collection (scraping). Corrections Reported as Completed but Not Verified in Production The following corrections have been documented as completed, but are not deployed on the VPS and therefore cannot be considered delivered until verified in production: Future-data leakage fix with permanent guard Deterministic and reproducible system (fixed seeds, stable ranking) Immutable IDs (combo_id) for each combination Protection against master file overwriting Hash validation (checksum) between generated and exported files Secure learning reset with archiving Diversity optimization (M6) Expanded historical data (7,268 draws) Rebuilt historical features (respecting format changes) What Needs to Be Done (Mandatory Scope) 1. Architecture Refactoring Split main.py (15,608 lines) into modular components: routers/ – HTTP endpoints services/ – business logic repositories/ – data access models/ – schemas and validation ml/ – Machine Learning and ranking pipeline 2. VPS Deployment Deploy all corrections, improvements, and expanded data to production environment. Configure the system to run stably and continuously. 3. Verification and Validation of Reported Corrections Validate that all listed corrections actually work in the production environment. Fix anything that is not working as expected. 4. Continuous Learning from the First Draw Configure the system to process and learn from the first available draw of each lottery. Respect chronological order and format/rule changes over time. Ensure learning is incremental and cumulative. 5. Learning Reset Button Implement (or verify and finalize) a button in the interface that: Deletes all previous learning Restarts processing from the first available draw Preserves immutable IDs and already generated master files Archives old data before deletion (safety) 6. Learning Evolution Progress Bar Implement (or verify and finalize) a visual indicator that shows: Historical processing progress (draws processed vs. total) Evolution of the stability metric between consecutive executions Charts and indicators of continuous learning 7. Master File Generation by Range Implement export of files by ID range. Preserve immutable IDs and original order (no renumbering). Allow download through the user interface. 8. Master File Maintenance with Hash Store each generated master file with its respective hash (checksum). Ensure the downloadable file is identical to the internally generated one. Provide file history for auditing. 9. Stability Metric Between Executions Calculate, store, and display the position difference of the first prize between one draw and the next. Display evolution on the dashboard with charts, alerts, and stability indicators. 10. Feedback Loop with Exponential Penalty Adjust the continuous learning mechanism to penalize large variations between consecutive executions. Apply exponential penalty when the difference exceeds the expected limit. 11. Incremental Reordering Replace full ranking reordering with local incremental adjustments. Preserve the relative position of the prize between executions, avoiding abrupt fluctuations. 12. Dashboard Refactoring Replace generic charts with actionable metrics: Evolution of the difference between consecutive executions Percentage of executions within expected limit Automatic alerts for critical variations Historical averages, medians, best and worst results Learning evolution progress bar 13. Pre-2005 Data Format Fix (El Gordo) Correctly handle draws prior to 2005 (format 6/49 vs 5/54). Ensure the system does not ignore or corrupt this data. 14. Daily Processing Automation Configure the system to run automatically after each new draw. Update rankings, metrics, and master files without manual intervention. Required Technical Skills Backend: Advanced Python (FastAPI, Pydantic, asyncio), modular code structuring. Database: MongoDB (pymongo), modeling and optimized queries. Frontend: React, TypeScript, Vite, REST API integration. Automation: Selenium, scraping, authenticated website navigation. Machine Learning: scikit-learn, PyTorch (existing models – no need to create new ones). Infrastructure: Linux, VPS, systemd, Git/GitHub. Plus: Docker, CI/CD, ranking optimization, immutable files, hash validation, stability metrics. Estimated Timeline The system has most of the code written, but nothing has been validated in production. Estimated timeline: 3 to 5 weeks, depending on the professional's experience. Payment Terms Payment per milestone, with validation of each stage before release. Milestones will be defined based on the scope above. Important Notes Source code already exists and is available. Many corrections have been reported as completed, but are not deployed on the VPS – therefore, they need to be verified, validated, and, if necessary, redone. No new Machine Learning models need to be developed. The focus is on finalization, organization, correct metrics, continuous learning, master file generation by range, verification of reported corrections, and deployment. The system must learn from the first available draw. Master files must be generated with immutable IDs, exported by range, and validated by hash. The reset button and learning evolution progress bar must be implemented or finalized and displayed on the dashboard. How to Apply Submit a proposal with: Brief presentation of your experience with the listed technical requirements. Suggested approach for refactoring, verification of reported corrections, continuous learning, master file generation by range, stability metrics, reset button, and learning evolution progress bar. Estimated timeline and detailed cost breakdown by milestones. Examples of previous work with similar systems (automation, ranking, dashboards, immutable files, stability metrics).
Skills required
This job is listed on Freelancer.com. AiZity aggregates listings for discovery only and is not the employer. To bid or apply, use the button in the sidebar.