LL-dc
log in

Advanced search

Message boards : Number crunching : LL-dc

Previous · 1 · 2 · 3 · 4 · 5 · Next
Author Message
Profile mikey
Avatar
Send message
Joined: 29 Apr 16
Posts: 63
Credit: 1,847,740,169
RAC: 18,374
Message 11824 - Posted: 3 Jun 2026, 15:12:54 UTC - in response to Message 11823.

I tried running it on my AMD AMD Radeon RX 9060 XT, I have 2 of them in the same pc and they both crunch most projects with no problems, but neither gpu returned valid tasks!!

Task 55787973
mikey ยท log out
Name LL-dc_85M_4_wu_171_8
Workunit 50225733
Created 1 Jun 2026, 21:55:53 UTC
Sent 2 Jun 2026, 2:26:27 UTC
Report deadline 9 Jun 2026, 2:26:27 UTC
Received 2 Jun 2026, 3:03:32 UTC
Server state Over
Outcome Computation error
Client state Compute error
Exit status 0 (0x0)
Computer ID 222552
Run time 4 sec
CPU time
Validate state Invalid
Credit 0.00
Device peak FLOPS 786.66 GFLOPS
Application version LL - dc v0.09 (opencl_ati_200)
Peak working set size 59.11 MB
Peak swap size 276.07 MB
Peak disk usage 2.65 MB
Stderr output
<core_client_version>8.2.9</core_client_version>
<![CDATA[
<stderr_txt>
2026-06-01 22:58:43 (11508): wrapper: running prpll.exe (-d 0 -v -use NO_ASM)
2026-06-01 22:58:43 (11508): wrapper: created child process 1200
20260601 22:58:43 PRPLL a3e4c7e starting
20260601 22:58:43 config: -d 0 -v -use NO_ASM
20260601 22:58:43 device 0, OpenCL 3661.0 (PAL,LC), unique id ''
20260601 22:58:43 85869361 No FFTs found in tune.txt that can handle 85869361. Consider tuning with -tune
20260601 22:58:44 85869361 config: -DNO_ASM=1
20260601 22:58:44 85869361 Using CARRY64
20260601 22:58:44 85869361 OpenCL: 3661.0 (PAL,LC):gfx1200, args -cl-finite-math-only -cl-std=CL2.0 -DNO_ASM=1 -DEXP=85869361u -DWIDTH=1024u -DSMALL_HEIGHT=256u -DMIDDLE=4u -DCARRY_LEN=8u -DNW=4u -DNH=4u -DAMDGPU=1 -DCARRY64=1 -DFFT_VARIANT=101u -DMAXBPW=4750u -DWEIGHT_STEP=0.038353673847609501 -DIWEIGHT_STEP=-0.036937004041686809 -DTAILT=U2(-7.52981578e-05f,0.0122715384f) -DTRIG_SCALE=9 -DTRIG_SIN={6.6579027251980952e-07,3.7209369054580932e-23,-4.9188217704570848e-20,1.0901995091303198e-33,-1.1506191102407305e-47,7.0839240359575376e-62,-2.8545227803597818e-76,8.0307778151820938e-91,} -DTRIG_COS={1,-2.2163834349100114e-13,8.1872592175725296e-27,-1.209740380477418e-40,9.5758876440886212e-55,-4.7164085661095887e-69,1.5838111582820031e-83,-3.830342691138796e-98,} -DTAILTGF31=U2(509684486u,293249438u) -DTAILTGF61=U2(249938719029223731ull,1245372627562045535ull) -DFFT_TYPE=4 -DWordSize=8u -DDISTGF31=819200u -DDISTWTRIGGF31=2560u -DDISTMTRIGGF31=768u -DDISTHTRIGGF31=131776u -DDISTGF61=1638400ull -DDISTWTRIGGF61=3072ull -DDISTMTRIGGF61=1792ull -DDISTHTRIGGF61=263040ull -DFRAC_BPW_HI=4061759487u -DFRAC_BPW_LO=4294967295u
20260601 22:58:44 85869361 FFT: 2M 4:1K:4:256:101 (40.95 bpw)
20260601 22:58:44 85869361 Loaded testTime : 328ms
20260601 22:58:44 85869361 V_NOP : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_ADD_I32 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_FMA_F32 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_ADD_F64 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_FMA_F64 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_MUL_F64 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 V_MAD_U64_U32 : 384307168202282304.00 cycles latency; time min: -1; avg 0
20260601 22:58:44 85869361 Loaded transposeIn : 323ms
20260601 22:58:45 85869361 Loaded readResidue -DREADRESIDUE=1: 305ms
20260601 22:58:45 85869361 LL loaded @ 0 : 0000000000000004
20260601 22:58:45 85869361 In file included from C:\Users\mike\AppData\Local\Temp\comgr-c7ff2d\input\CompileSource:1:
In file included from C:\Users\mike\AppData\Local\Temp\comgr-c7ff2d\include\fftp.cl:5:
In file included from C:\Users\mike\AppData\Local\Temp\comgr-c7ff2d\include\fftwidth.cl:18:
C:\Users\mike\AppData\Local\Temp\comgr-c7ff2d\include\fftbase.cl:1136:8: error: assigning to '__private F2' (aka '__private float2') from incompatible type 'ulong2' (vector of 2 'ulong' values)
1136 | base = U2(fma(a, -w.y, w.x), fma(a, w.x, -w.y));
| ^ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
Error: Failed to compile source (from CL or HIP source to LLVM IR).

20260601 22:58:45 85869361 Compiling 'fftp.cl' error COMPILE_PROGRAM_FAILURE (-15) (args -cl-finite-math-only -cl-std=CL2.0 -DNO_ASM=1 -DEXP=85869361u -DWIDTH=1024u -DSMALL_HEIGHT=256u -DMIDDLE=4u -DCARRY_LEN=8u -DNW=4u -DNH=4u -DAMDGPU=1 -DCARRY64=1 -DFFT_VARIANT=101u -DMAXBPW=4750u -DWEIGHT_STEP=0.038353673847609501 -DIWEIGHT_STEP=-0.036937004041686809 -DTAILT=U2(-7.52981578e-05f,0.0122715384f) -DTRIG_SCALE=9 -DTRIG_SIN={6.6579027251980952e-07,3.7209369054580932e-23,-4.9188217704570848e-20,1.0901995091303198e-33,-1.1506191102407305e-47,7.0839240359575376e-62,-2.8545227803597818e-76,8.0307778151820938e-91,} -DTRIG_COS={1,-2.2163834349100114e-13,8.1872592175725296e-27,-1.209740380477418e-40,9.5758876440886212e-55,-4.7164085661095887e-69,1.5838111582820031e-83,-3.830342691138796e-98,} -DTAILTGF31=U2(509684486u,293249438u) -DTAILTGF61=U2(249938719029223731ull,1245372627562045535ull) -DFFT_TYPE=4 -DWordSize=8u -DDISTGF31=819200u -DDISTWTRIGGF31=2560u -DDISTMTRIGGF31=768u -DDISTHTRIGGF31=131776u -DDISTGF61=1638400ull -DDISTWTRIGGF61=3072ull -DDISTMTRIGGF61=1792ull -DDISTHTRIGGF61=263040ull -DFRAC_BPW_HI=4061759487u -DFRAC_BPW_LO=4294967295u -DFFT_FP64=0 -DFFT_FP32=1 -DNTT_GF31=1 -DNTT_GF61=1 )
20260601 22:58:45 85869361 Can't compile fftp.cl
20260601 22:58:45 Exception "Can't compile fftp.cl"
20260601 22:58:45 Bye
2026-06-01 22:58:45 (11508): prpll.exe exited; CPU time 0.000000
2026-06-01 22:58:45 (11508): called boinc_finish(0)

</stderr_txt>
<message>
upload failure: <file_xfer_error>
<file_name>LL-dc_85M_4_wu_171_8_0</file_name>
<error_code>-240 (stat() failed)</error_code>
</file_xfer_error>
</message>
]]>

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11825 - Posted: 3 Jun 2026, 15:35:20 UTC - in response to Message 11824.

We had the same issue with a 9070 on win while linux was running, reported.

Profile [DPC] hansR
Send message
Joined: 9 Jun 15
Posts: 17
Credit: 1,805,821,190
RAC: 463,456
Message 11827 - Posted: 4 Jun 2026, 7:37:24 UTC

After running multiple jobs on my RTX 5080, the task
https://srbase.my-firewall.org/sr5/result.php?resultid=56125592 errored out

last part of the log:

20260603 19:36:32 85772207 83000000 e57869b18d185a2d 295.0 ETA 00:14 20260603 19:41:28 85772207 84000000 672e36804ef9677e 295.4 ETA 00:09 20260603 19:46:22 85772207 85000000 68cc6b05b5cacf67 294.5 ETA 00:04 20260603 19:48:23 Exception gpu_error: OUT_OF_RESOURCES (-5) clGetEventInfo(event, CL_EVENT_COMMAND_EXECUTION_STATUS, sizeof(status), &status, 0) at src/clwrap.cpp:415 getEventInfo 20260603 19:48:24 Bye 2026-06-03 19:48:24 (27556): prpll.exe exited; CPU time 3989.593750 2026-06-03 19:48:24 (27556): called boinc_finish(0) </stderr_txt> <message> upload failure: <file_xfer_error> <file_name>LL-dc_85M_3_wu_176_5_0</file_name> <error_code>-240 (stat() failed)</error_code> </file_xfer_error> </message> ]]>

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11828 - Posted: 4 Jun 2026, 8:46:47 UTC - in response to Message 11827.

Interesting, the tasks was nearly done but errored out. I will report it but their forum is down at the moment. Credits are granted for this WU.

Drago75
Send message
Joined: 29 Mar 21
Posts: 10
Credit: 72,538,832
RAC: 617,678
Message 11832 - Posted: 5 Jun 2026, 22:06:10 UTC

The task that I mentioned earlier with the number 55785824 finally completed after around 15 hours run time on my RTX 3070-ti but strangely logged only 1h28. I had to pause it twice for computer shutdown. The Stderr file started logging the last 3,5 hours and also shows the last interuption but everything before that is gone. So the official runtime is 1h28 instead of 15h.

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11833 - Posted: 6 Jun 2026, 8:33:15 UTC - in response to Message 11832.
Last modified: 6 Jun 2026, 8:37:21 UTC

as mentioned in the FAQ. I did a GPU time request to the dev to save it at the end near CPU time.

tito
Send message
Joined: 30 Dec 14
Posts: 20
Credit: 1,245,663,063
RAC: 1,558,117
Message 11838 - Posted: 7 Jun 2026, 16:48:18 UTC - in response to Message 11833.

Host https://srbase.my-firewall.org/sr5/show_host_detail.php?hostid=251426
1080ti returns only errors. For now I switch back to TF

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11839 - Posted: 7 Jun 2026, 17:05:01 UTC - in response to Message 11838.

reported, nearly the same issue from the other user.

ahorek's team
Send message
Joined: 5 May 23
Posts: 6
Credit: 4,570,067
RAC: 72,870
Message 11840 - Posted: 7 Jun 2026, 21:22:13 UTC

I have the same issue as mikey on linux & windows

In file included from /tmp/comgr-9915-19-4657d4/include/fftp.cl:5:
In file included from /tmp/comgr-9915-19-4657d4/include/fftwidth.cl:18:
/tmp/comgr-9915-19-4657d4/include/fftbase.cl:1136:8: error: assigning to '__private F2' (aka '__private float2') from incompatible type 'ulong2' (vector of 2 'ulong' values)
1136 | base = U2(fma(a, -w.y, w.x), fma(a, w.x, -w.y));

Mr P Hucker
Avatar
Send message
Joined: 30 Sep 17
Posts: 38
Credit: 48,022,324
RAC: 4,017
Message 11841 - Posted: 8 Jun 2026, 1:03:32 UTC

Running on a Radeon R9 290, 100% according to Boinc after only 15 hours, stderr.txt as below, but still producing heat. I assume it's on another section? I shall leave it running. Unfortunately the incorrect progress means Boinc scheduler left it for a bit until it panicked.

12:22:25 (15388): wrapper (7.7.26016): starting 12:22:25 (15388): wrapper: running gp.exe (spt.txt) GP/PARI CALCULATOR Version 2.17.0 (released) amd64 running mingw (x86-64/GMP-6.1.2 kernel) 64-bit version compiled: Sep 28 2024, gcc version 12-posix (GCC) threading engine: single (readline v8.0 enabled, extended help enabled) Copyright (C) 2000-2024 The PARI Group PARI/GP is free software, covered by the GNU General Public License, and comes WITHOUT ANY WARRANTY WHATSOEVER. Type ? for help, \q to quit. Type ?18 for how to get moral (and possibly technical) support. parisize = 8000000, primelimit = 1048576, factorlimit = 1048576 **********begingp************* 9917938322621 9917938322623 end_of_algo **********closegp************* Goodbye! closegp************* 20:12:33 (15388): gp.exe exited; CPU time 14594.187500 20:12:33 (15388): called boinc_finish(0)


P.S. I love the "type ?18 for moral support"!

Stef42
Send message
Joined: 22 Dec 14
Posts: 26
Credit: 486,440,540
RAC: 7,785
Message 11844 - Posted: 8 Jun 2026, 17:05:58 UTC
Last modified: 8 Jun 2026, 17:34:49 UTC

Hi Rebirther! Is there an option (somehow) to unpack the app (prpll-linux64-v3n), run a tune session (creates a config.txt + tune.txt file) and repack?

Or can I use app_config.xml to point to a cuda version of prpll?
I have no idea whether the project accepts this or not.

Why do I want this you might ask?! On ll-dc runs (85M), a tune session that takes just ~5-10 minutes improves the speed by +30%. Runtime decreases from 12:40 (HH:MM) to 08:00.

You might not be aware of all functions of prpll, but generally a tune can improve the speed of a run by a lot.

Curious what you thoughts are!

EDIT: also the hardcoded "NO_ASM" is a major slowdown on nvidia-cards. (like -20%). I can see why it might help older cards (or AMD?) but modern nvidia-cards suffer quite a lot from this switch.

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11845 - Posted: 8 Jun 2026, 17:20:23 UTC - in response to Message 11844.

I never tested -tune but in this case run a standalone session with -tune after unpacking the file. There should be a config file after that. A cuda version doesn't help because opencl is fast as cuda.

If you have more infos let me know.

Stef42
Send message
Joined: 22 Dec 14
Posts: 26
Credit: 486,440,540
RAC: 7,785
Message 11846 - Posted: 8 Jun 2026, 17:37:52 UTC - in response to Message 11845.
Last modified: 8 Jun 2026, 17:39:41 UTC

I figured an easier way. Copied the package from /var/lib/boinc/, ran tune and pasted the line from config.txt to a app_config.xml file.

That works well, already improved by 16-20%. The only thing in the way is the hardcoded "NO_ASM", which I cannot get rid off. I can understand why its there, but its a major slowdown on modern nvidia cards (4xxx,5xxx).

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11847 - Posted: 8 Jun 2026, 17:49:22 UTC - in response to Message 11846.

yes, the dev's solution doesn't work without -NO_ASM for 1xxx cards. How much percent do you loose with this option? We can get rid off if there is a newer working version.

Stef42
Send message
Joined: 22 Dec 14
Posts: 26
Credit: 486,440,540
RAC: 7,785
Message 11848 - Posted: 8 Jun 2026, 17:57:42 UTC - in response to Message 11847.
Last modified: 8 Jun 2026, 17:59:26 UTC

Did some more checking, actually its a lot less than I thought. About 4.5% lost due to "NO_ASM".

In a standalone version (which has config.txt and tune.txt) the iteration time is 346 and ETA is roughly 08:00 (HH:MM)

Vanilla boinc app (opencl_nvidia_200) ETA is roughly 12:40 (HH:MM)

Using app_config (copy line from config.txt) I can get the iteration time down to 450 and ETA is roughly 10:26.

Pausing boinc and pasting tune.txt in slot 0 and restarting again, I can get the iteration time down to 364 and ETA is roughly 08:23.

Ofcourse this is not useful to do every time, but it demonstrates how much time we leave on the table. My situation: 4070 Ti NVIDIA card.

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11849 - Posted: 8 Jun 2026, 18:02:26 UTC - in response to Message 11848.

I will contact the dev to change something in code.

If you have an app_config file you only need to change the extended option once.

zombie67 [MM]
Avatar
Send message
Joined: 4 Dec 14
Posts: 33
Credit: 2,061,370,984
RAC: 10,074,815
Message 11852 - Posted: 10 Jun 2026, 19:48:00 UTC
Last modified: 10 Jun 2026, 19:48:15 UTC

GPU max WUs in progress limit = 1


Can this be increased to 2? My 5090 is running at only about 80% load.

Profile rebirther
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist
Avatar
Send message
Joined: 2 Jan 13
Posts: 8533
Credit: 199,516,813
RAC: 7
Message 11853 - Posted: 10 Jun 2026, 19:54:10 UTC - in response to Message 11852.

no, 1 is enough for better scaling.

Profile [AF>Libristes]Maeda
Send message
Joined: 6 Jun 19
Posts: 5
Credit: 42,827,755
RAC: 2,510
Message 11854 - Posted: 10 Jun 2026, 20:27:27 UTC - in response to Message 11810.

I am trying a work unit on a NVidia GTX 750 Ti, 3 days already and it shows ETA 9 days. That will be a little bit far from its deadline, isn't it? Still worth to keep?


Just wondering that is running but I have extended the deadline for this WU.

Finally completed in success in 12 days.

zombie67 [MM]
Avatar
Send message
Joined: 4 Dec 14
Posts: 33
Credit: 2,061,370,984
RAC: 10,074,815
Message 11855 - Posted: 11 Jun 2026, 1:08:34 UTC - in response to Message 11853.

no, 1 is enough for better scaling.

I am not sure what this means exactly. For example, if I can run 1 task in 5 hours, or two at a time in 9 hours, I should be running two at a time to maximize output. I don't want 20% of my GPU's capacity wasted. I would prefer to run 2x of the LL-dc. But if I can't run two LL-dc tasks, I will be forced to run something else (like einsteinium or PG) as the second task. That just slows down the LL-dc task, without any increase in LL-dc over-all output.

Previous · 1 · 2 · 3 · 4 · 5 · Next
Post to thread

Message boards : Number crunching : LL-dc


Main page · Your account · Message boards


Copyright © 2014-2026 BOINC Confederation / rebirther