Message boards : Cafe Rosetta : Other projects.
Previous · 1 . . . 4 · 5 · 6 · 7
| Author | Message |
|---|---|
|
Sid Celery Send message Joined: 11 Feb 08 Posts: 2604 Credit: 47,220,881 RAC: 0 |
An old status update, but significant, so I post it as a marker WCG Status Update 2026-08-21 Updates on the status of the Africa Rainfall Project (ARP1), Mapping Arthritis Markers (MAM1, beta30), and Mapping Cancer Markers (MCM1) projects at WCG. Updates on web and API development that is currently in progress will be provided next week. ARP1 - ARP1 workunits are being distributed again, consider the project intermittent pending a full restart. ARP1 workunits are orders of magnitude larger than MAM1 and MCM1 workunits, so the in-memory cache across a partitioned BOINC backend that allowed multiple backend servers to share the load, and work around the slower NFS performance and database IO at Nibi compared to the dedicated appliance we had at Graham, is simply less beneficial for ARP1. Further, it negatively impacts MCM1 and MAM1 for little benefit, as we simply cannot hold enough ARP1 workunits resident in memory to make a dent in throughput anyways. We now have a working ARP1 lifecycle based on our NFS and block storage as we await access to object storage which should increase throughput down the road, and we will continue to send limited bursts of ARP1 work to establish what the consistent rate will be in the fall. - We await final confirmation from TU Delft that the thousands of results we have validated and shipped to them from the new BOINC backend so far are consistent with earlier results and we should keep them coming. We have no reason to believe otherwise, but this must be confirmed explicitly and we expect movement on this by mid-September, 2026, if not before. Until then we will continue to distribute workunits and update volunteers. The project remains scientifically and technically relevant based on discussions with TU Delf, and we are exploring plans to expand climate research at WCG using the ARP1 model. - The ARP1 specific files generations.txt, state.txt, and completed.txt will be published again when the tasks mentioned above are further along. We anticipate mid-September, after we review with TU Delft, as the earliest reasonable date. We could restart the script as is, and it would run against the legacy database as the stats dump currently does, but we would rather restructure it to connect directly to the new BOINC database while producing files in the exact same format so we can avoid more synchronization issues. - ARP1 extremes from older generations were reviewed, some restarted successfully, some still have issues that must be resolved before they can be distributed. Extremes and BOINC client errors that have now been masked by the transitioner as having too many errors to continue distributing resends are being assessed. When we are able to restore the ARP1 specific stats export as mentioned in the previous bullet point, progress or lack thereof on these will become visible as before. MAM1 - We have begun testing a ROCm AMD GPU build for MAM1 on Linux in the beta30 project. - Updated versions of the Windows and Linux NVIDIA CUDA GPU builds and multi-threaded CPU builds for the Mapping Arthritis Markers (MAM1) beta30 project have been released, and will be promoted to production MAM1 next week. We plan to release and maintain Windows, Mac, Linux, Android on AMD, Intel, Arm, NVIDIA, Apple Silicon and Metal over time. - We have since made significant changes to MAM1 over the summer of 2026, improving the neural network architecture, changing memory management on the GPU to provide a guard against OOM crashes on volunteer devices, changing the optimizer from Simulated Annealing to Differential Evolution to fix the low GPU utilization which previously remained even after changing the cross validation logic to train multiple folds simultaneously. Other bugs as they were discovered or reported on the forum were fixed, and there will likely be more as we continue the rollout. - All recently issued work under the beta30 and MAM1 project has been viable science for some time. Gene signatures are evaluated between batches to plan and configure subsequent batches. We have implemented and tested an automated procedure for gathering promising genes and subsets of genes from the data lake where signatures from previous batches are stored, and minting new batches based on the best performing signatures we have. We will continue to build up this infrastructure with the expectation it will be applied to MCM1 and other diseases going forward. For this reason, occasional Ovarian and Lung batches have been released and will continue to be released as MAM1 beta30 work. MAM1 gives us MCM1 on GPU, which we plan to pursue soon. - We are working to resolve packaging issues for the DLLs published with some of our Windows releases for beta30, which has recurred multiple times as we iterate on the beta30 application versions. Although we are able to fix these problems with subsequent releases each time, this has happened multiple times in the beta30 release lineage. Some beta30 applications were also promoted to a MAM1 application version before we identified the issue with the beta30 bundle. The issue continues to affect some volunteers, often showing up as runtime errors, and we are working on a more formal release process to ensure this doesn't happen again. - We introduced a regression to multi-threaded CPU-only MAM1, now fixed. OpenMP and LibTorch should now respect BOINC --nthreads again, thank you to volunteers who reported this in the forum. Appears to have been caused by an added include of torch.h in a CPP source file where it should not have been. MCM1 - MCM1 beta31 project that runs on NVIDIA/AMD GPUs will be added, soon. However, we must do this in a way that allows currently participating devices to continue to have an impact as they process MCM1, while benefiting from the increased capabilities of the MAM1 GPU capable platform run on the MCM1 datasets. What we determine here will also be the path toward allowing less capable devices to run MAM1 productively. Likely, the backends other than LibTorch, dlib and OpenCV, will be used to train and evaluate SVM and RandomForest classifiers, with the same OpenMP based heuristic search between signature or signature population evaluations provided by our serializable build of the ensmallen library. While this will require a bit of development to ensure those backends still behave well with all the changes we have made while prioritizing the LibTorch backend, these workunits would run on older machines and support followup investigation or rapid hypothesis testing for the most promising gene signatures first identified by MCM1 with LibTorch. While the neural network architecture developed for MAM1 already works for MCM1, there must still be a separate analysis and development phase before MCM1 on GPU is ready to release within the MCM1 application lineage. However, the foundation has been laid with MAM1 application development, and the implementation of the data lake supported by the postgres DuckDB plugin ecosystem. - Validation for new work MCM1 work is finally healthy as volunteers have reported in the forum, but we continue to recover pending validations.
|
[VENETO] bobovizSend message Joined: 1 Dec 05 Posts: 2215 Credit: 13,720,774 RAC: 0 |
An old status update, but significant, so I post it as a marker Where did you find these info? P.s. in user's profile there is no possibility to select gpu calculation, up to now We wait.... |
Message boards :
Cafe Rosetta :
Other projects.
©2026 University of Washington
https://www.bakerlab.org