First known occurrence:
https://dev.azure.com/dnceng/public/_build/results?buildId=462402&view=results
Commit info:
HEAD is now at 7dc30fc1 Merge 03693d335194dd76177cec10ac159f32de973082 into 0a29c61468f60f815119e52e477d0154351a4abf
Proximate failure:
C:\dotnetbuild\work\100ccf8f-adbf-49ff-87b2-37c7a2ef03ce\Work\c74808db-757a-43c3-a1f2-ec1c841e07ff\Exec>C:\dotnetbuild\work\100ccf8f-adbf-49ff-87b2-37c7a2ef03ce\Payload\CoreRun.exe C:\dotnetbuild\work\100ccf8f-adbf-49ff-87b2-37c7a2ef03ce\Payload\xunit.console.dll JIT\Stress\JIT.Stress.XUnitWrapper.dll -parallel collections -nocolor -noshadow -xml testResults.xml -trait TestGroup=JIT.Stress C:\dotnetbuild\work\100ccf8f-adbf-49ff-87b2-37c7a2ef03ce\Work\c74808db-757a-43c3-a1f2-ec1c841e07ff\Exec>set _commandExitCode=-1073741701
C000007B = STATUS_INVALID_IMAGE_FORMAT
As this happens on ARM, a usual cause is erroneously mixing up ARM and Intel binaries.
/cc @dagood
@jkotas can this be closed now you reverted the commit?
We have been dealing with several different CI breaks. The revert fixed the failures in runtime-live-build (Runtime Build Windows_NT x86 release).
This issue tracks failures in CoreCLR Pri0 Test Run Windows_NT arm. It is not fixed yet. I was not able to pin point it to any single commit. I believe that it was likely introduced by some kind of ambient infrastructure change.
I experimentally tried to roll back two commits merged in between the last passing and the first failing run that seemed potentially related (Leandro's change to assembly resolution logic and JanV's merge of Viktor's build cleanup) but, as JanK rightly pointed out, sadly neither fixed the regression.
Today I tried a different approach: I checked out a private branch off the "old coreclr" repo master, I backported Jarret's change re-enabling ARM runs and I ran an "old coreclr PR pipeline" on the change:
Apparently it failed with exactly the same failure as the new runs. I believe this confirms that the failure is more likely related to the ambient state of the ARM machines than to "real" pipeline changes, just as JanK speculated in his previous response. It's probably time to start talking to infra guys, initially perhaps @ilyas1974, @MattGal or @Chrisboh, but I doubt anyone will be able to shed more light on this before the end of the holidays.
Where there specific questions you had about the ARM hardware being used? The current hardware being used for Windows ARM64 runs is running a newer build of Windows that we were previously using as well as Python 3.7.0 - the x86 version.
Thanks so much Ilya for your quick response in the holiday period :-). As a first question, I guess I'm wondering what happened between 2019/12/19 10:45 AM and 12:14 PM PST w.r.t. Windows ARM pool so that, without any relevant commits to the repo in that time interval (there are four commits there total, I can provide more detailed info if needed), a run at 10:45 AM passed and a subsequent one at 12:14 PST failed.
Just for reference, I believe that the name of the Helix queue in question is "Windows.10.Arm64.Open".
That was around the time period we migrated from the systems running the hold Seattle chipset to the Centrix chip set. We also migrated from using python 2 to python 3. We also update to a newer build of Windows. The old systems are still around (just disabled in Helix).
Great, looks like we're starting to untangle the problem. Have you got any idea whether we might be able to somehow divide the problem space to identify which of the three components of the update caused the regression?
I have a system (we're still the process of updating the rest of the systems in that queue) in the windows.10.arm64 queue with the same software configuration, but different hardware. If we are able to run these workloads on that system, that will tell us if it's hardware specific or if it's something with the software configuration.
Awesome, just please let me know what to alter in the queue definitions or whatnot and I can easily trigger a private run to evaluate that.
I have disabled all systems in the windows.10.arm64 queue except for the one with the latest build of Windows and Python 3. If you send your workloads to that queue, it will be run against this system.
Hmm, apparently I've messed something up in the experimental run
Sending Job to Windows.10.Arm64... F:\workspace\_work\1\s\.packages\microsoft.dotnet.helix.sdk\5.0.0-beta.19617.1\tools\Microsoft.DotNet.Helix.Sdk.MonoQueue.targets(47,5): error : ArgumentException: Unknown QueueId. Check that authentication is used and user is in correct groups. [F:\workspace\_work\1\s\src\coreclr\tests\helixpublishwitharcade.proj]
Please let me know what I need to change when you have a chance and I'll retry the run.
Thanks a lot
Tomas
I talked with my team and it appears that it is not as simple a task to run your workloads on the Windows.10.Arm64 queue as I first thought. Why don't we do this, let's coordinate a good time for me to move the laptop to the Windows.10.Arm64.Open queue and we can make it the only available system in that queue. What time works best for you on this?
The centriq machines cannot run arm32. The arm32 workloads we have should target the the queue where the laptops end up in, and we should stop running arm32 windows until that queue is up and running.
/cc @trylek @ilyas1974
We're still waiting for the new laptop queue to be deployed into production and MLS to complete the deployment of the devices.
Then the change https://github.com/dotnet/runtime/pull/1283 @sandreenko has in flight should go in quickly. Follow up question:
Where are the seattle machines currently, still in the windows.arm64 queue? If so can we move them to a new windows arm32 queue (which will eventually be where the laptops end up)
@ilyas1974 - I'm back in the office now so I'm ready to carry out any experiments as needed once you let me know that we have a candidate queue suitable for testing. Thank you!
@trylek @ilyas1974 It looks like Windows ARM jobs are still disabled in innerloop testing "runtime-coreclr" pipeline.
Is there any progress in fixing this?
The systems we have that will function for both ARM64 and ARM32 Windows based workloads are still being provisioned by DDFUN/MSIT. We do have dedicated ARM32 systems that may meet your needs - they can be found in the windows.10.arm32.open queue.
In the issue https://github.com/dotnet/core-eng/issues/8490 @jashook expressed concerns that the queue windows.10.arm32.open is underpowered for CoreCLR - have we got any idea how bad the situation is and / or whether we might be able to use this queue at least temporarily, before the new machines become available? I find it somewhat scary that Windows ARM testing has been down for a month by now.
We originally had Raspberry PI systems in that queue. The systems that are now in this queue are Hummingboards which are more powerful and what Windows uses for their ARM32 workloads. Is it possible for you to give these systems a try and see how they work for your needs?
My pleasure :-).
@ilyas1974 - I have run an experimental PR sending ARM32 work to the queue Windows.10.Arm32.Open as you suggested but the "Send to Helix" step seems to be failing with
The response contained an invalid status code 404 Not Found
@trylek I was able to submit work to it just now without issue, I'll look at the linked PR and see if I have suggestions.
@trylek I see no evidence in the linked PR runs of any attempts to send work to windows.10.arm32.open and I had no problem sending some myself (2ecd1d23-7b9c-461b-a3aa-3cdba5fef409)
So, if you need help you'll need to provide some details / logs of timeframe, how you sent the work, etc.
Poking around (the yml has changed a lot recently so I may be mistaken / incomplete here) I think you might just need to uncomment this line too: https://github.com/dotnet/runtime/blob/c05a44c4feadfc04cd2e4dae8c44e9e3e4b49947/eng/pipelines/runtime.yml#L136 to test it out.
Thanks @MattGal for the investigation. My primary indication that things should work stemmed from the pipeline log that used to work before just fine:
Uploading payloads for Job on Windows.10.Arm32.Open... Finished uploading payloads for Job on Windows.10.Arm32.Open... Sending Job to Windows.10.Arm32.Open... Uploading payloads for Job on Windows.10.Arm32.Open... Finished uploading payloads for Job on Windows.10.Arm32.Open... Sending Job to Windows.10.Arm32.Open... Sent Helix Job 45bf7ee0-051a-4cbe-9d8a-2f29d2dde243 Sent Helix Job f5988133-c8d8-46c4-8a37-7c137d52a2ab Waiting for completion of job 45bf7ee0-051a-4cbe-9d8a-2f29d2dde243 Waiting for completion of job f5988133-c8d8-46c4-8a37-7c137d52a2ab Job f5988133-c8d8-46c4-8a37-7c137d52a2ab is completed with 25 finished work items.
Can it be the case that the queue name is now case-sensitive so that I messed up by Pascal-casing it? I'm not sure what other means I know to decisively claim that we're sending work to the right queue but I'll be happy to add any instrumentation per your suggestion to the pipelines to understand this better. Sadly the "runtime.yml" pipeline you pointed out is different - while in theory it's also worth fixing, in the smoke test I ran (#1696) the runtime-coreclr pipeline is not using it in any manner.
Thanks for your patience!
Tomas
Queue names aren't case sensitive (though some logging will preserve the case used), but if you have logs of the 404 happening I can try to dig deeper.
Hmm, I'm afraid that the only logs I have right now are from the smoke test I ran: https://dev.azure.com/dnceng/public/_build/results?buildId=481937&view=logs&jobId=41021207-15b4-5953-02cc-987654ff0f7b&j=41021207-15b4-5953-02cc-987654ff0f7b&t=786f87fb-d4be-5ad9-2d0b-e89af519eba6
As I said, I can easily run another test with added logging as necessary to let us better understand what's happening.
Hmm, I'm afraid that the only logs I have right now are from the smoke test I ran: https://dev.azure.com/dnceng/public/_build/results?buildId=481937&view=logs&jobId=41021207-15b4-5953-02cc-987654ff0f7b&j=41021207-15b4-5953-02cc-987654ff0f7b&t=786f87fb-d4be-5ad9-2d0b-e89af519eba6
As I said, I can easily run another test with added logging as necessary to let us better understand what's happening.
That's what I was looking for, thanks. I'll take a look now.
I retried the failing legs on the smoke test run out of curiosity whether your e-mail regarding health of some machines in the Windows.10.Arm32.Open Helix queue might indicate improved behavior of the queue but sadly I received the 404 error codes again:
The problem is pretty clear, and I think support @jashook 's previous assertions unfortunately.
Basically, from the logs there's OOMs and stack overflows, but the 404s (sample log) are caused by not running the work item at all, thus the console log API 404'ing.
Log message:
2020-01-15 13:30:32,793: INFO: executor(306): _download_to: Signal response processing thread to terminate as it is running for more than 10 mins
2020-01-15 13:30:32,799: INFO: executor(439): _get_unpacked_file_paths: Reading archive file:\\?\C:\data\helix\work\89638daa-12fe-45e7-93b6-3fd24415f7ac\Work\092e9096-1d75-45b1-8ea7-3c6b85963ade\ce13c6a9-438a-4977-a0bd-9bfe02f10374.zip.partial
2020-0
So, we can explore why the network is extra slow , we can talk about allowing downloads to take > 10 minutes, but for now these machines aren't going to be good for your purposes.
(@ilyas1974 can you loop someone with network expertise about these machines?)
John has the most experience here and can look at the switch level to see what is going on. @JpratherMS, can you please take a look?
@JpratherMS @ilyas1974 @MattGal If the Windows ARM machine problem is a network issue, has anyone investigated the network slowness issue?
@trylek and all: what are the next steps to get Windows ARM and Windows ARM64 jobs running again? Just troubleshoot the networking issues? Use different machines?
@BruceForstall - I worked with Matt on running a couple of experimental runs against the queues he suggested but none worked so far due to various infra issues. I believe that, once we receive a queue we're able to run our pipelines on, that should be basically it. Please just note that most of this discussion dealt with 32-bit ARM - in fact, I haven't yet investigated whether we might be already able to re-enable the ARM64 runs using the Windows.10.Arm64.Open queue that kind of initially started this thread due to no longer supporting ARM32 - I can easily look into that.
/cc @tommcdon for visibility
@trylek What's the current status of Windows arm32/arm64 testing?
@BruceForstall - I re-enabled Windows.Arm64 PR job last Monday, see e.g. here:
[I grabbed a PR that happened to succeed to make a good impression ;-).] For Windows arm32, there seem to be a new queue of the Galaxy Book machines DDFUN is standing up, @ilyas1974 sent out an update earlier today, I'll send an initial canary run to their new queue and I guess next week we can decide on re-enabling the Windows arm32 PR jobs if it turns out to be stable enough.
There are currently 18 Galaxy book systems online with this new queue
Nice!
Arm32 is now being run again
arm32 library test runs are still disabled against this issue: https://github.com/dotnet/runtime/blob/5092e107acc5e9895a35d0cb17b9a25f15f46a63/eng/pipelines/runtime.yml#L945. Is that intentional? cc @safern
The comment should be removed; we've removed almost all Windows arm32 testing: https://github.com/dotnet/runtime/pull/39655
So we don't run our libraries tests against arm32 anywhere at all anymore?
Specifically Windows arm32; Linux arm32 testing should be everywhere.
I don't know about the libraries. For coreclr, I think we still do builds (but not test runs) in CI, and test runs only in the "runtime-coreclr outerloop" pipeline (but none of the many stress pipelines).
For libraries I believe we agreed to no longer test against arm32. We think that the chances of finding an arm32 bug on libraries tests is very low, however JIT does need to do testing so we scaled that testing to outerloop pipeline.
chances of finding an arm32 bug on libraries tests is very low,
I think it was @jkotas observation - I believe it, but it is not impossible and it feels odd to offer first class support without ever running libraries tests, even once a cycle. I am not going to push on this if consensus is it's not necessary/feasible. Do we get coverage from dotnet/iot repo?
chances of finding an arm32 bug on libraries tests is very low,
I have made this comment in the context of doing arm32 tests for every PR.
I agree that we should have a once a day or once a week test runs of the full matrix for everything we ship as officially supported.
@jaredpar thoughts about this? we should have no tests on Windows ARM32, and occasional but regular runs (that include regular library unit tests) on Linux ARM32.
I believe the reason why we stopped running tests was because win-arm32 is not a supported platform on .NET 5. https://github.com/dotnet/core/blob/master/release-notes/5.0/5.0-supported-os.md
Actually, should stop building the win-arm32 runtime pack?
We should have stopped all Windows ARM32 activities so far as I can see.