I have a long-running .NET console application (4.6.1). Suddenly since fews days I am getting this error in Visual Studio,
Unhandled exception at 0x77256214 (ntdll.dll) in My.exe: 0xC0000374: A heap has been corrupted (parameters: 0x77272378).
Pressing F5 again, it shows me another stuck and then stuck with the same error,
Exception thrown at 0x771B234D (ntdll.dll) in My.exe: 0xC0000005: Access violation reading location 0x656C6573.
I have global try catch which is not triggering in the console application. I have added,
<legacyCorruptedStateExceptionsPolicy enabled="true" /> and [HandleProcessCorruptedStateExceptionsAttribute] but still application crashes without hitting the breakpoint. Then I tried WinDBG (.loadby sos clr then !analyze -v then !CLRStack -a then !dumpstackobjects),
OS Thread Id: 0x3da4 (0)
Child SP IP Call Site
007cea38 771ac33c [GCFrame: 007cea38]
007cea54 771ac33c [HelperMethodFrame_1OBJ: 007cea54] System.Threading.Monitor.ReliableEnter(System.Object, Boolean ByRef)
007cead0 60f98710 System.Collections.Hashtable+SyncHashtable.Remove(System.Object) [f:\dd\ndp\clr\src\BCL\system\collections\hashtable.cs @ 1518]
007ceafc 0db7405f Oracle.DataAccess.Client.OracleResourcePool.RemoveResourceHolder(Oracle.DataAccess.Client.OracleResourceHolder)
007ceb28 0db73ffd Oracle.DataAccess.Client.OracleResourceHolder.Dispose()
007ceb34 0db7386f Oracle.DataAccess.Client.OracleResourceHolder.TransactionCompleted(System.Object, System.Transactions.TransactionEventArgs)
007ceb38 0086e053 [InlinedCallFrame: 007ceb38]
007cebb0 0086e053 [MulticastFrame: 007cebb0] System.Transactions.TransactionCompletedEventHandler.Invoke(System.Object, System.Transactions.TransactionEventArgs)
007cebdc 556cc0fb System.Transactions.InternalTransaction.FireCompletion()
007cebf0 556ea3f2 System.Transactions.TransactionStatePromotedAborted.EnterState(System.Transactions.InternalTransaction)
007cec08 556ed9c0 System.Transactions.TransactionStateDelegatedAborting.ChangeStatePromotedAborted(System.Transactions.InternalTransaction)
007cec14 556ee8ba System.Transactions.DurableEnlistmentDelegated.Aborted(System.Transactions.InternalEnlistment, System.Exception)
007cec24 556e34d6 System.Transactions.SinglePhaseEnlistment.Aborted()
007cec6c 0e36d662 Oracle.DataAccess.Client.PromotableTxnMgr.Rollback(System.Transactions.SinglePhaseEnlistment)
007cec90 556ed922 System.Transactions.TransactionStateDelegatedAborting.EnterState(System.Transactions.InternalTransaction)
007cecd4 556eb8f5 System.Transactions.TransactionStateDelegated.Rollback(System.Transactions.InternalTransaction, System.Exception)
007cece4 556c9c0a System.Transactions.Transaction.Rollback()
007ced18 556f0f43 System.Transactions.TransactionScope.InternalDispose()
007ced50 556f0da7 System.Transactions.TransactionScope.Dispose()
007cee04 0d49f9ae MyApp.Program.method()
Clearly, it's related to TransactionScope and Oracle but dunno why its crashing. For now I am just looking for a work around
maybe creating a debugger dump can help as I don't think the issue has enough data to look at.
@tarekgh I can send you the in your (can't share it publicly) whats your email?
@imranbaloch my email is already listed in my profile https://github.com/tarekgh.
@tarekgh Thanks Tariq I shared the file with you. Any help will be appreciated.
I got the dump file but I am seeing some problem with it:
0:000> dx Debugger.Sessions[0].Processes[34732].Threads[15780].Stack.Frames[1].SwitchTo();dv /t /v
Debugger.Sessions[0].Processes[34732].Threads[15780].Stack.Frames[1].SwitchTo()
@edi void * hHandle = 0x00000170
007ce8b8 unsigned long dwMilliseconds = 0xffffffff
007ce8bc int bAlertable = 0n0
007ce88c long Status = 0n0
007ce888 union _LARGE_INTEGER * pTimeOut = 0x00000000
007ce880 union _LARGE_INTEGER TimeOut = {-4294966928}
007ce85c struct _RTL_CALLER_ALLOCATED_ACTIVATION_CONTEXT_STACK_FRAME_EXTENDED Frame = struct _RTL_CALLER_ALLOCATED_ACTIVATION_CONTEXT_STACK_FRAME_EXTENDED
0:000> !handle 0x00000170
ERROR: !handle: extension exception 0x80004002.
"Unable to read handle information"
Maybe you need to create a full dump with the handle information.
Also, the failure here looks a little bit strange because it says
Break instruction exception - code 80000003
while I am not seeing any break instruction on the place it stopped.
eax=00000001 ebx=00b098a8 ecx=00b09af0 edx=0be9dda7 esi=00000000 edi=00000170
eip=771ac33c esp=007ce83c ebp=007ce8ac iopl=0 nv up ei pl nz na pe nc
cs=0023 ss=002b ds=002b es=002b fs=0053 gs=002b efl=00000206
ntdll!NtWaitForSingleObject+0xc:
771ac33c c20c00 ret 0Ch
are you hitting this on any other machine? this could be something wrong with this specific machine?
@tarekgh I created dump from Task Manager when the crash happened. Yes it happens on 2 different server machines.
This is obvious heap corruption problem and need to have more live investigation. Most properly some code corrupted allocated memory (or even stack memory). The best you can do is to debug the app using one of the following mechanisms to know who is corrupting the memory:
https://techcommunity.microsoft.com/t5/IIS-Support-Blog/Debugging-Heap-corruption-with-Application-Verifier-and/ba-p/376844
http://www.daviddahlbacka.com/BugCleaner/DebuggingHeapCorruption.pdf
I am going to close this issue as no action can be done from our side. Feel free to reply back if you think we can help more here.
Just for info its not from our code. Its started from System.Transactions.TransactionScope.Dispose. We are using class like,
static void Main(string[] args)
{
using(TransactionScope scope = new TransactionScope())
{
// Our Code
// No exception till here
}// exception starts here during Dispose
}
That's why we are helpless in this case. The problem is that exception cannot be catch. We are thinking to migrate to a language other than .NET because this case makes our business people too much unrealistic that's why We are trying our best to stick with the same code base (even we can try .NET core if the issue can be fixed after porting).
Any way thanks a lot for helping and investigating. Any other help will be appreciated
@HongGit @jimcarley do you have any suggestion for @imranbaloch regarding TransactionScope.Dispose?
@jkotas can you think in a way can help with this issue.
@imranbaloch maybe it is still worth trying to investigate the heap corruption as I mentioned before. It is possible your code corrupted some memory and the side effect just appeared during the TransactionScope.Dispose.
Is it possible you create a sample repro?
As you know, when the code leaves the "using" statement block for the TransactionScope, the TransactionScope.Dispose method is invoked.
If the code inside the "using" block has not yet called "scope.Complete()", the TransactionScope.Dispose method assumes that not all of the code inside the "using" block successfully executed, so the transaction must be aborted. Typically the "scope.Complete()" statement would be the last statement before the ending brace "}" of the "using" statement block, like this:
using(TransactionScope scope = new TransactionScope())
{
// Our Code
scope.Complete();
}
Since the scope was not completed, it forces the abort/rollback of the transaction and then also attempts to notify any delegates for the Transaction.TransactionCompleted event to let them know that the outcome of the transaction is "aborted". That is why you see the System.Transactions.InternalTransaction.FireCompletion method on the stack.
The Oracle code must have registered a delegate for the Transaction.TransactionCompleted event using the method Oracle.DataAccess.Client.OracleResourceHolder.TransactionCompleted as the delegate.
The heap corruption is apparently detected during the processing in the Oracle code.
I am wondering if the "using" using block actually encountered some exception that prevented it from calling "scope.Complete()", but that exception is being masked by another exception that occurs during the TransactionCompleted processing which is happening "inline" in this case.
Maybe if you put a try..catch(Exception) inside the "using" block, with the "catch" just before the end of the "using" block. Then you might catch the "original" exception (if my theory is correct). The catch block should then rethrow the exception. Like this:
using(TransactionScope scope = new TransactionScope())
{
try
{
// Our Code
scope.Complete();
}
catch (Exception ex)
{
// Do something you might want to do.
throw;
}
}
I don't have any insight into what that Oracle code is doing other than what is on the stack. I think further investigation needs to start with Oracle. If it comes back that indicates that System.Transactions may have corrupted something, then please re-raise to our attention.
@tarekgh unfortunately no because the code we have talk with internal databases which are very sensitive for business (Full Disclosure I work for government entity in UAE and we are Microsoft enterprise client and we get a lot of support from Microsoft support tickets). We have ported our code from VB6 (some time ago) where we never got heap corrupted exception . But as you suggested some links (thanks for sharing), we will continue our investigation.
@jimcarley we debug the code it actually not calling the scope complete method(means it was trying to rollback). There was no exception just when it execute the the last brace, the process crashes. I have open a case with Oracle (thanks for suggestion) https://community.oracle.com/message/15427797#15427797
But I am wondering why process crashing even though I have <legacyCorruptedStateExceptionsPolicy enabled="true" /> and [HandleProcessCorruptedStateExceptionsAttribute].
Also don't you guys think there should be some flag that make the developer to catch these type of exceptions. In the current package manager era, every application using a bunch of third party libraries and not able to catch an exception from these third party libraries is disappointing (at least for me).
We also tried to use Oledb with System.Transactions but found lot of issues that's why we choose ODP.NET.
Anyway, I appreciate all of your help.
I am not able to answer your questions about the process crashing issue, but will continue to try to address the System.Transactions questions.
What I gather from your last message is that your "using (TransactionScope)" block does, in fact, include a "scope.Complete()" call, but that call is not executing. And you are not seeing an exception? Or you are not able catch the exception? The TransactionScope.Dispose will execute even if there was an exception thrown from within the "using" block. I believe because the process crashing is due to something in the Oracle code during the processing of TransactionCompleted event which is invoked during TransactionScope.Dispose, you don't get to see what exception is happening inside the TransactionScope using block. That is why I suggested the try..catch within the using block. That catch should execute BEFORE TransactionScope.Dispose is invoked, so you should be able to see the "original" exception. Of course, maybe that is less interesting to you than why the Oracle code is causing the process to crash.
@jimcarley thanks for your suggestions.
And you are not seeing an exception?
During debugging, we found exception but we are catching and logging this.
Or you are not able catch the exception?
Yes, we are unable to catch the heap corruption exception.
What you suggest, we already have in place in our code. I just wanna sketch our high-level code base and tell you what exactly happened,
{
// Get some data from database
foreach (var item in items)
{
try
{
using (var scope = new TransactionScope())
{
var commit = true;
try
{
// Some Code
}
catch (Exception ex1)
{
//log exception
commit = false;
}
if (commit)
{
scope.Complete();
}
}
}
catch (Exception ex)
{
//log exception
}
}
}
catch (Exception ex2)
{
//log exception
}
What we found during debugging is that commit=false and then process crashed. There are daily more than 10k items processed and it just happens for some items. For now, we are removing items that lead to process crash but it's very strange for us that it happens just for some items not all. Note we have a lot of other items that have commit=false but no crash happens.
@imranbaloch,
What exception are you seeing for "ex1", where you set "commit" to false?
Does that exception type vary when you see the process crash? Or is there some sort of pattern that might indicate a different between the crashing and non-crashing scenarios?
I am not sure we are going to get to the bottom of this until we get some information from your request to Oracle, as that's where the crash is happening.
Jim
What exception are you seeing for "ex1", where you set "commit" to false?
When I update a record in Oracle, I see an exception from some trigger to stop this update.
Sometimes, an exception from a trigger causes the process to crash but sometimes it doesn't crash. Honestly, I don't see a specific pattern but for sure whenever the process crashes there is an exception from a trigger.
Thanks, Jim.
Okay. Sounds like we need input from Oracle now.
Have a good weekend.
Jim
Most helpful comment
As you know, when the code leaves the "using" statement block for the TransactionScope, the TransactionScope.Dispose method is invoked.
If the code inside the "using" block has not yet called "scope.Complete()", the TransactionScope.Dispose method assumes that not all of the code inside the "using" block successfully executed, so the transaction must be aborted. Typically the "scope.Complete()" statement would be the last statement before the ending brace "}" of the "using" statement block, like this:
using(TransactionScope scope = new TransactionScope())
{
// Our Code
scope.Complete();
}
Since the scope was not completed, it forces the abort/rollback of the transaction and then also attempts to notify any delegates for the Transaction.TransactionCompleted event to let them know that the outcome of the transaction is "aborted". That is why you see the System.Transactions.InternalTransaction.FireCompletion method on the stack.
The Oracle code must have registered a delegate for the Transaction.TransactionCompleted event using the method Oracle.DataAccess.Client.OracleResourceHolder.TransactionCompleted as the delegate.
The heap corruption is apparently detected during the processing in the Oracle code.
I am wondering if the "using" using block actually encountered some exception that prevented it from calling "scope.Complete()", but that exception is being masked by another exception that occurs during the TransactionCompleted processing which is happening "inline" in this case.
Maybe if you put a try..catch(Exception) inside the "using" block, with the "catch" just before the end of the "using" block. Then you might catch the "original" exception (if my theory is correct). The catch block should then rethrow the exception. Like this:
using(TransactionScope scope = new TransactionScope())
{
try
{
// Our Code
scope.Complete();
}
catch (Exception ex)
{
// Do something you might want to do.
throw;
}
}
I don't have any insight into what that Oracle code is doing other than what is on the stack. I think further investigation needs to start with Oracle. If it comes back that indicates that System.Transactions may have corrupted something, then please re-raise to our attention.