We recently moved to exadata. In a day at least 3 times, users cannot login in. It shows signing in. There is no crash report.
Hi Steve,
Thank you to everyone for all your help.
We finally fixed the issue. It is quite wierd. One of the connection pools in the rpd was not checked for Enable Connection Pooling. We just tried that. That fixed the issue. Possible exadata was keeping all those sessions and not releasing them. Did not happen in earlier db versions.
Really, really appreciate all the help.
Thanks,
Paami.
@paamikumar
From logs I can see that your nqsserver was crashing till May 6 also I can see some query cancelled After 48 Minutes of running and also you mentioned you switched to exadata DB and issue start to happen adding all these facts together I can conclude that
your nqsserver may get into hang status from time to time which lead to this issue I will recommend the following to narrow down the issue
1- When issue happen run the following Query in your DB Side
select count(*),program from gv$session where LAST_CALL_ET > 600 and status = 'INACTIVE' group by program;
If you getting a big Number then I will recommend to edit NQSConfig.INI File you can control the query Time through the following Two Variables
DEFAULT_DB_MAX_EXEC_TIME = 600;
MAX_LOGICAL_QUERY_EXEC_TIME = 600;
NQSServer need to be restarted so changes can take affect
Values above can be changed based on your estimation how long the longest query run
We have been able to pin point the issue. Sessions are created and not cleared. When we open the rpd online, there are more than 2000 sessions. Once we clear the sessions, we are able to login.
Issue still continuing. Any suggestions where we need tweaking.
You need to figure out where those sessions come from. A sane implementation doesn't break even when 2000 users are logged on and I doubt you have 2k concurrent users who are causing this.
What is keeping those connections open? Where do they even start their lifecycle? What queries are being forced where that cause this?
Hi Christian,
When we open the rpd online, we see these session of users which last used = Never.
Even after we purged these sessions, we had to restart the obis1 and obips1.
It is just not making any sense to us as well.
Attached is a screenshot.
Paami
Those are cache entries. Not sessions.
We see these cache entries. When we run the command netstat -a | grep -i established | wc -l, the sessions are increasing to > 2000.
I checked the Service Request. There are a lot of "moving parts" for investigation. I see discussions on OAM SSO, Load Runner tests, etc. There appear to be some outstanding action plans that should be completed. 1. Update to the latest patched client tools 2. Clean diagnostic dump after the issue reproduces to assist in narrowing the issue/ check for any crashes, etc. 3. Questions about agent schedule/backlog 4. RPD configuration for Exadata Oracle DB source needs some updates. 5. Checking for initialization block failures, hangs
6. Checking ulimits. Previously, a 'broken pipe' was reported, which I see in the logs attached here.. that can indicate the nqsserver crashed (as Moustafa) mentioned, but there is no .out file attached here to correlate that.
Next time you restart, notate the PID of the obipsX, obisX, when the issue occurs, check if the PID changes (that is a easy clear indication of crash), in addition there should be crash reports generated in the component logs directory. Once the issue is narrowed, the thread can be updated; otherwise, it may be guess after guess here without proper data to perform root cause analysis. The service request is also 24*7, which can be helpful in some cases, but also not efficient if there is no traction on the issue, or new lines of thinking without follow through. I would suggest you align it with someone in your timezone. The assigned person has a team they can collaborate with, but this really needs a thorough methodical review.
BTW, you should change your netstat command should be filtered by the nqsserver port (default: :9514, but you can check with status.sh -v command)
status.sh
@paamikumar In the obis diagnostic log, i can see init block failures cause by invalid SQL statement. Please validate the failed queries and then check the signin issues. Also make sure you have selected Correct Database Type When Designing An RPD Refer:- (Doc ID 2965721.1)
There is an SR open, all the above has been suggested already. There is progress being made, all will be summarized post solution.
Description: Join the OCI Enterprise AI Product Management team for a deep dive into the latest product roadmap updates and upcoming innovations. During this webinar, we’ll cover key roadmap highlights, showcase product demos, and share insights into what’s ahead for OCI Enterprise AI. The session will conclude with a live…
Hi! I am looking for audit logging for Analytics objects, such as users creating, modifying, or deleting datasets, dataflows, schedules, etc. The traditional Presentation Services component provides such logging in domains/oas/servers/obips1/logs/auditjson.log, which contains audit information for catalog objects,…
I'm trying to index a Local Subject Area so I can use it in my AI Agent, however, I keep receiving the generic error below. Request status(500) : error = {"prefix":"DSS","code":50000,"message":"There was an unexpected error while processing this request. Please try again."}
it is observed that in OAC while using case statements the output is not expected.For example if we have a case statment like case when table1.column1=table2.column1 then 'yes' else 'No' end as Attribute1 Now if i use this attribute with a fact-amount column attribute 1|| Amount on a report then we dont get the aggregation…
Once I select that value from the list box Region I created, I want to filter the table by my selection. See the images