<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: FunCorp</title>
    <description>The latest articles on DEV Community by FunCorp (@funcorp).</description>
    <link>https://dev.to/funcorp</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F879754%2F74e9cfdb-e2bd-43e0-a280-fe3a2be76659.jpg</url>
      <title>DEV Community: FunCorp</title>
      <link>https://dev.to/funcorp</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/funcorp"/>
    <language>en</language>
    <item>
      <title>Stopping Zombie Objects Invasion at iOS</title>
      <dc:creator>FunCorp</dc:creator>
      <pubDate>Thu, 11 Aug 2022 08:53:51 +0000</pubDate>
      <link>https://dev.to/funcorp/stopping-zombie-objects-invasion-at-ios-2o99</link>
      <guid>https://dev.to/funcorp/stopping-zombie-objects-invasion-at-ios-2o99</guid>
      <description>&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--Rjw70VbM--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/f1nngrcsynjszpn27yz1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--Rjw70VbM--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/f1nngrcsynjszpn27yz1.jpg" alt="Image description" width="880" height="495"&gt;&lt;/a&gt;&lt;br&gt;
So you turn on your laptop, open Crashlytics, and voila: EXC_BAD_ACCESS objc_release… What now?&lt;/p&gt;
&lt;h2&gt;
  
  
  Let’s discuss the nature of such crashes
&lt;/h2&gt;

&lt;p&gt;All NSObject heirs and Pure Swift objects are reference types. When passing a new variable to an object, we copy only the memory location address in the heap, not the entire contents.&lt;/p&gt;

&lt;p&gt;When object reference semantics are handled incorrectly, we may unknowingly cause a memory leak. It can be a trivial retain cycle between two objects referring to each other.&lt;/p&gt;

&lt;p&gt;In order to avoid leaks, Objective-C and Swift have special reference options: weak and unowned. While weak is one of the good guys, you should be extremely careful with unowned!&lt;/p&gt;

&lt;p&gt;Firstly, it is possible to miscommunicate with an unowned link, even within a single thread. Secondly, the heap is shared by all application threads, so non-atomic access can be unsafe. Thirdly, in release builds, the compiler can convert our unowned(safe) references to unowned(unsafe) type during optimization. These are faster but more dangerous.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Let’s see what can happen when using an unowned(unsafe) reference leading to an object that has already been destroyed, or partially destroyed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In the first case, the address may contain unprepared raw memory. Then, you’ll get the EXC_BAD_ACCESS KERN_INVALID_ADDRESS error.&lt;/p&gt;

&lt;p&gt;Now, if there is a new object at the address, maybe even of the same type, you may not experience any crashes. However, the problem will persist and get accumulated in the app. This is clearly demonstrated by the following snippet:&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;With any luck, we’ll see the coveted “bar” in the console. And if there’s none, we’ll get a crash with the EXC_BAD_ACCESS error.&lt;/p&gt;

&lt;p&gt;These errors are called “use after free”. Such issues can even lead to exploits in your apps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Apple says
&lt;/h2&gt;

&lt;p&gt;The bowels of the CoreFoundation framework contain NSZombie, an excellent mechanism. The idea behind it is simple: when starting a process, there is a check whether the “NSZombieEnabled” environment variable is present. If it’s found, the magic begins.&lt;/p&gt;

&lt;p&gt;For NSObject heirs, the dealloc method is redefined, and the object class is substituted by NSZombie (by isa ivar via object_setClass) via the runtime library. The memory is not freed after the substitution, and a leak occurs. Any access to the object triggers assert, which reports the object type and the name of the method invoked.&lt;/p&gt;

&lt;p&gt;Then, we enable debugging, activate the environment variable and do &lt;a href="https://developer.apple.com/documentation/xcode/investigating-crashes-for-zombie-objects"&gt;various manipulations&lt;/a&gt; in the app.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Naturally, like many others, we began looking for zombie objects in this way.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;iFunny is a product that earns money through advertising, and we cooperate with a dozen partner SDKs. &lt;a href="https://medium.com/@FunCorp/what-ios-developers-should-be-prepared-for-when-integrating-in-app-advertising-in-2022-7b15016c37ce"&gt;Dealing with memory leaks&lt;/a&gt; and similar issues is an integral part of an iOS engineer’s routine in products like this.&lt;/p&gt;

&lt;p&gt;However, the official approach has several shortcomings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Manually searching for zombie objects is not much fun at all.&lt;/li&gt;
&lt;li&gt;There is no way to activate the mechanism in the prod for users.&lt;/li&gt;
&lt;li&gt;And even if it were possible to activate it on the prod, what is to be done with endless memory leaks?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This means we need another option.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Mike Ash says
&lt;/h2&gt;

&lt;p&gt;While fishing for information online, we came across an article. It describes a custom Zombie mechanism implementation using the public API of the Objective-C runtime library.&lt;/p&gt;

&lt;p&gt;We scoured the open-source solutions and found several ready-made implementations. Here they are:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/lilidan/NSZombie"&gt;https://github.com/lilidan/NSZombie&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chromium.googlesource.com/chromium/src/+/179d013b254fce69feb811badfad8cc0cd4952f2/components/crash/core/common/objc_zombie.mm"&gt;https://chromium.googlesource.com/chromium/src/+/179d013b254fce69feb811badfad8cc0cd4952f2/components/crash/core/common/objc_zombie.mm&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Dokay/DJZombieCheck"&gt;https://github.com/Dokay/DJZombieCheck&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/AlexTing0/DDZombieMonitor"&gt;https://github.com/AlexTing0/DDZombieMonitor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;They inspired us to create an in-house zombie mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  Funcorp’s way
&lt;/h2&gt;

&lt;p&gt;Just like we do with other modules of our app, we implemented this solution as an SPM package. Let’s look at its structure:&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;The package consists of two targets. We pull the Swift.fatalError function, sorely lacking in Objective-C, from “swift_shims”. Under the hood, the transmitted message is logged in all the necessary system stuff (stderr, sys logs , etc.), and abort is invoked.&lt;/p&gt;

&lt;p&gt;Now, the FNZombie main target is where the real magic happens. It contains Objective-C code and some tight deallocs/work with the runtime library, so you’ve got to resort to the -fno-objc-arc (compile without ARC) compilation flag. We tried putting some functions that can’t be used in ARC code into a separate target and pull them as a dependency; unfortunately, it wouldn’t work that way.&lt;/p&gt;

&lt;p&gt;Publicly, there’s only one header with a single interface sticking out of the target:&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;We use it to turn the zombie object search mechanism on and off in the main app.&lt;/p&gt;

&lt;p&gt;In order to implement the zombie object mechanism, a root class has to be created. FNZombie.h file&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;At the very least, a root class must contain a static initialize method, which is a runtime requirement.&lt;/p&gt;

&lt;p&gt;Next, we define a forwardingTargetForSelector instance &lt;a href="https://developer.apple.com/documentation/objectivec/nsobject/1418855-forwardingtargetforselector"&gt;method&lt;/a&gt;. If the object can’t define the method being invoked, then the message sent is processed here. This is the perfect place to cause a fatalError.&lt;/p&gt;

&lt;p&gt;We also define a bunch of methods that can be invoked while sending the message; you can also slip in a fatalError there.&lt;/p&gt;

&lt;p&gt;After engaging our mechanism, we allocate a buffer where we will store and run our zombie objects. Overriding the dealloc method for objects:&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;The zombie_dealloc substituted method is a normal C function.&lt;/p&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;We’ve got to do a few things inside:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieve the object class and size&lt;/li&gt;
&lt;/ol&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;Release all variables of the instance and “associated objects”&lt;/li&gt;
&lt;/ol&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;“Zero” the memory in the object&lt;/li&gt;
&lt;/ol&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;Substitute the object class for zombie using the isa ivar&lt;/li&gt;
&lt;/ol&gt;


&lt;div class="ltag_gist-liquid-tag"&gt;
  
&lt;/div&gt;


&lt;p&gt;And finally, put our zombie in the buffer!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/funcorp/FNZombie"&gt;The implementation is available here&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Obviously, this mechanism is dangerous and unstable. We decided that we will only detect zombies for beta users in Test Flight. We also highly discourage you from using this solution in production.&lt;/p&gt;

&lt;p&gt;A couple of releases later, we got the coveted message in Crashlytics:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--UrxTRR05--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/tpz38s3vitjqsvfqmi46.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--UrxTRR05--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/tpz38s3vitjqsvfqmi46.png" alt="Image description" width="880" height="185"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This made us happy campers, as we pinpointed the issue. After getting rid of it, we continued to monitor use after free problems.&lt;/p&gt;

</description>
      <category>ios</category>
      <category>mobile</category>
      <category>objc</category>
      <category>swift</category>
    </item>
    <item>
      <title>Practical Guide to Create a Two-Layered Recommendation System</title>
      <dc:creator>FunCorp</dc:creator>
      <pubDate>Fri, 22 Jul 2022 08:24:16 +0000</pubDate>
      <link>https://dev.to/funcorp/practical-guide-to-create-a-two-layered-recommendation-system-2oal</link>
      <guid>https://dev.to/funcorp/practical-guide-to-create-a-two-layered-recommendation-system-2oal</guid>
      <description>&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--ZbvUzzpF--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/llu7ls76vnp6jo1e70j6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--ZbvUzzpF--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/llu7ls76vnp6jo1e70j6.png" alt="Image description" width="700" height="364"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Written by: Gleb Abroskin&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At iFunny, we are trying to compose the best possible feed of memes and funny videos. To rate our job, people use smile/dislike buttons and comments. Some of them even post memes about our efforts.&lt;/p&gt;

&lt;p&gt;The idea of why we need multiple layers in RecSys is already described here. This article will look at how to deploy such a system and talk about a high-level architectural overview of the recommendation services at FUNCORP.&lt;/p&gt;

&lt;p&gt;First of all, let’s define what we want to achieve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Useful definitions
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Online feature — a feature, the value of which must be updated as soon as possible, but no later than one minute.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Offline feature — a feature, the value of which might be updated once per hour/day/week.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Feature group — the list of features in the specified order.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model features — are the list of feature groups used to construct the input tensor for the model.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A/B shuffle — generation of A/B tests, used for shuffling users to the new groups if an experiment significantly impacted vital metrics.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Requirements
&lt;/h2&gt;

&lt;p&gt;The system must:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Support multiple projects.&lt;/li&gt;
&lt;li&gt;Support multiple ongoing A/B tests meaning that — models might be trained on the subset of the data selected by users of the A/B group; datasets might be prepared concerning users’ A/B group and current shuffle.&lt;/li&gt;
&lt;li&gt;Support up to 50 deployed models, with the maximum size of a single model in memory up to 5GB.&lt;/li&gt;
&lt;li&gt;Support different models, such as rank models and candidate selection models.&lt;/li&gt;
&lt;li&gt;Be able to update models as frequently as once an hour. “Update model” in this context means to deploy a retrained version of the same model.&lt;/li&gt;
&lt;li&gt;Support both online and offline features for inference and training.&lt;/li&gt;
&lt;li&gt;Reduce resource utilization at least in half. I must note here that our old solution consisted of 25 instances of quite pricey AWS servers running 24/7.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now let’s design a system that is going to satisfy our requirements&lt;/p&gt;

&lt;p&gt;Theoretically, we would need five components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt; Service with an HTTP API will provide inference models (inference service).&lt;/li&gt;
&lt;li&gt;Service to schedule data preparations, training, and model deployments.&lt;/li&gt;
&lt;li&gt;Database or another repository to store features for training and inference.&lt;/li&gt;
&lt;li&gt;Service to run feature computations (Spark, Kafka Streams, Flink, etc.).&lt;/li&gt;
&lt;li&gt;External data sources: Kafka with actions, Database with user/content data.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We would need two databases for different workloads and a cache for inference service.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--zaBCZfXm--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/q6ey9b8yaafemw3q78a6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--zaBCZfXm--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/q6ey9b8yaafemw3q78a6.png" alt="Image description" width="700" height="364"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Data sources
&lt;/h2&gt;

&lt;p&gt;There is not much to say about data sources. I will just give an approximation of the data amount we are dealing with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;12B+ events in Kafka are being read both online and in batches.&lt;/li&gt;
&lt;li&gt;Hundreds of thousands of content entries monthly.&lt;/li&gt;
&lt;li&gt;Tens of millions of monthly active users.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Inference service
&lt;/h2&gt;

&lt;p&gt;The service is relatively straightforward but has some interesting details. For example, it is written in Kotlin as the rest of our backend. We decided to dump Python mainly because of developer productivity and worse tooling than Kotlin/Java. To make Kotlin work for ML model inference, we convert all models to ONNX format and premise them with ONNX runtime (Java). By doing this, we can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Simplify deduction of any model to a single function call.&lt;/li&gt;
&lt;li&gt;Avoid having multiple ML runtimes in the container.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because ONNX allows us to be model-agnostic, we can replace or add a model in production with a single HTTP request.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--QZLMcnC6--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/fybb0xor4dbxfyhscakl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--QZLMcnC6--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/fybb0xor4dbxfyhscakl.png" alt="Image description" width="700" height="294"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As you might see, we have two different resources: “…/first-tier-model/” and “…/second-tier-model/.” They represent “layers” or “tiers” in the RecSys architecture. The resources have a lot of parameters that are caused by all types of things: the A/B-tests model is being used for test groups, projects, etc.&lt;/p&gt;

&lt;p&gt;The downside of using ONNX is that we have to convert every model deployed to production, but the process is automated and takes only about 50 lines of Python code.&lt;/p&gt;

&lt;p&gt;Inference service consists of three instances of the exact HTTP. How do we make sure they use the same model for every layer/experiment and other configuration options? We went with an event sourcing—like the approach here, the instance that receives a command to update the model, writes config in MongoDB, and then reads it from “change stream.” All the updates must be done via reading the change stream, not serving the HTTP.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scheduler
&lt;/h2&gt;

&lt;p&gt;We went with Airflow because it checks all the boxes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Allows to schedule with different intervals, start times, timeouts, etc.&lt;/li&gt;
&lt;li&gt;Allows backfills.&lt;/li&gt;
&lt;li&gt;Supports streaming workloads (tasks that must work 24/7) deployment via “@once.”&lt;/li&gt;
&lt;li&gt;Allows deploying pods to K8s.&lt;/li&gt;
&lt;li&gt;Supports native spark + K8s integration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We use all those features and start to write our provider with needed hooks, operators, and other goodies, for example, an operator to get current A/B-tests generation.&lt;/p&gt;

&lt;p&gt;Another thing Airflow does for us is CD: once the model is trained, the Airflow task deploys the model to ML-inference via HTTP API.&lt;/p&gt;

&lt;h2&gt;
  
  
  “Feature store” (FS)
&lt;/h2&gt;

&lt;p&gt;Instead of deploying or using existing feature stores, we went with just two databases to avoid the organizational load coming with FS. Our databases of choice are ClickHouse for storing historical data used in model training and analytics and MongoDB for online stores. Mongo was chosen because we have expertise in deploying and maintaining significant clusters of it. However, we use it basically as a key-value store, where a key is user/content ID and value is a feature group.&lt;/p&gt;

&lt;p&gt;Both databases might store raw values instead of final feature values. For example, if the model needs smile rate (#smile/#views) database would contain both numbers of smiles and the number of views. This workaround was used to optimize DB load in streaming applications because without calculating the rate it is possible to write increments in DB without reading the latest value from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Features computation
&lt;/h2&gt;

&lt;p&gt;We use Spark to do all kinds of data tasks. Here are just a few examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured streaming tasks to compute the “user actions” feature with 20s latency.&lt;/li&gt;
&lt;li&gt;Batch tasks to prepare training data for NFM models.&lt;/li&gt;
&lt;li&gt;Data validation tasks containing business logic, for example: “X Kafka topic can not be empty,” “Distribution of column Y in ClickHouse must match the distribution of field Z in Mongo.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is A LOT to talk about Spark, especially streaming, and our experience with doing everything in Kotlin. There will be another article about it. Stay tuned :).&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployment
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--e5MXXyRb--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/jdm0uq2fanuikw8y7bry.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--e5MXXyRb--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/jdm0uq2fanuikw8y7bry.png" alt="Image description" width="700" height="364"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Оur Airflow instance is deployed as a standalone service alongside other services in our data center. The host with Airflow also contains Java and Spark 3.2.1 installation to be able to deploy Spark apps natively in K8s.&lt;/p&gt;

&lt;p&gt;Our K8s clusters contain multiple nodes in our DC and various nodes in AWS, which allows us to optimize the costs while having an opportunity to extend resources if needed.&lt;/p&gt;

&lt;p&gt;For now, our workload consists of KubernetesPodOperator and SparkSubmitOperator. Still, we are migrating to Airflow on K8s, which allows us to write Python code instead of wrapping it in a container and using KubernetesPodOperator.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--DS2VwrQr--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/e13lc713uori27dnzu46.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--DS2VwrQr--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/uploads/articles/e13lc713uori27dnzu46.png" alt="Image description" width="700" height="563"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s add a simple online feature — “the number of user interactions” — to the system, e.g., likes and comments grouped by content type.&lt;/p&gt;

&lt;p&gt;To calculate those features, we would use the stream of user events from Kafka and content type from MainMongo. Joining will be done by Mongo Spark connector, but I should note here that we could not use native join since it loads the whole collection from Mongo to the memory. Instead, we are making batch queries for the data collected in the span of Spark’s processingTime.&lt;/p&gt;

&lt;p&gt;To deploy the application, we are going to write an Airflow DAG containing all the configurations for:&lt;/p&gt;

&lt;p&gt;a) Online job.&lt;/p&gt;

&lt;p&gt;b) Offline job.&lt;/p&gt;

&lt;p&gt;c) Online import.&lt;/p&gt;

&lt;p&gt;The DAGs consist of a single SparkSubmitOperator, which launches spark applications in the cluster mode inside the K8s cluster. After the initial import online job starts to work every five seconds and imports all the events to the online store while features are being read and cached by the ML service. The same Spark application with different run parameters writes features to the offline store used for training ML models. By sharing the code, we minimize the possibility of an error and reduce the time to market for the features.&lt;/p&gt;

&lt;p&gt;Airflow also serves us as a CD tool: once the model is trained and validated, we use HTTP API to update the model in production. This was made possible by using ONNX for all the models in the inference service since they have the same interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Future plans
&lt;/h2&gt;

&lt;p&gt;New architecture met all the requirements. By implementing it, we were able to achieve set goals within the quarter of the year, which I would consider a success.&lt;/p&gt;

&lt;p&gt;This is what we plan to do next:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Trying Reinforcement Learning in the recommendations. There is a wonderful presentation by Oskar Stål from Spotify on why RL is the future of RecSys. Our architecture will allow us to store users’ cohorts and their changes in the training process while using the current cohort in the inference service and update it on the fly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Service to control data skews and better monitoring overall. We have set up Spark’s metrics collection from the standalone and K8s clusters. Still, there is a lot of work left with assessing the quality of the content which goes into RecSys training, making sure that online jobs can process all the events and checking feature values against DWH as a source of truth, etc.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Exactly-once guarantees for the data transformations. We went to production with the at-least-once warranty provided by Spark Structured Streaming with checkpoints enabled. For a single guarantee, we plan to implement CAS (check if the batch was written in the database and apply changes in the transaction only if it was not) using BatchID provided by foreachBatch function. There are a couple of edge cases. For example, what to do if a long BatchID overflows? But all of them are solvable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Migrate completely to spark-kotlin-API or Scala, adapting pySpark. It’s easier for the DS team to use pySpark API, and although it will add some operational load, we are considering adding it to the stack. In terms of Kotlin, everything works fine, and even streaming support was added, but we are still not quite sure which language will be the main in our repo.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Better CI/CD pipeline. There are automation scripts for everything needed to deploy a new feature group. Still, we are launching them manually since we are trying to figure out release schedules and other operation patterns. Once we figure out the process, we are going to automate everything.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scaling up the number of models. At some point, the inference service won’t be able to host all the needed models for all projects and experiences. We understood this risk from the beginning and decided that when the time comes, we will implement client-side service discovery while continuing to update inference service configurations from Airflow.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Hope we gave you more insights as to why we use multiple layers in RecSys and how we deploy such a system. Lastly, we hope we gave a deeper dive on how we utilize a high-level architectural overview of the recommendation services at FUNCORP.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>11 Traits of a Senior QA Engineer</title>
      <dc:creator>FunCorp</dc:creator>
      <pubDate>Wed, 22 Jun 2022 09:51:18 +0000</pubDate>
      <link>https://dev.to/funcorp/11-traits-of-a-senior-qa-engineer-3ekb</link>
      <guid>https://dev.to/funcorp/11-traits-of-a-senior-qa-engineer-3ekb</guid>
      <description>&lt;p&gt;Written by: Alexey Anisimov, head of QA at &lt;a href="https://medium.com/@funcorp"&gt;Funcorp&lt;/a&gt; &lt;/p&gt;

&lt;p&gt;There is a tremendous variety of QA Engineer vacancies, ranging from junior to lead tester and even to principal QA Engineer. We’re often asked what qualities a senior-level tester should have compared to junior or middle-level ones. Let’s try to answer this.&lt;/p&gt;

&lt;p&gt;Our Head of QA identified a set of qualities and behaviors required for a true Senior QA Engineer. Here, he shares these observations with you.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Take responsibility
&lt;/h2&gt;

&lt;p&gt;There is a goal to reach. For example, the rollout of a new microservice or moving a UI button from one place to another. Senior QA Engineers will take responsibility for the result, communicate with specialists, find those who can help, solve the challenges along the way, and do much more. In short, they will do anything to accomplish the task. And they do not need to have everything explained.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Work it out all the way to the end
&lt;/h2&gt;

&lt;p&gt;Sometimes, when you work on a problem, you get to the point of not having enough knowledge, such as when it is unclear why a new service does not create a queue in Kafka or why the application crashes when you leave the home screen.&lt;/p&gt;

&lt;p&gt;In my opinion, the experienced Senior QA Engineers will not shift responsibility to specialists with sufficient knowledge (“let the developers/DevOps/product managers figure it out”). They will continue to be responsible and dig deeper by engaging other people as needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Choose the right tools
&lt;/h2&gt;

&lt;p&gt;Often, I hear something like “I won’t write in Java,” “I prefer Python,” “I will only use Cypress,” “only Appium,” and so on. Senior QA Engineers will examine and apply the most effective tool to address the specific task.&lt;/p&gt;

&lt;p&gt;They don’t have universal “silver bullets” but select their tools based on input requirements. For example, if they want to organize the automation of testing, they will choose the technology and languages that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Are suitable for automating a specific product&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Are easily supported&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Have a candidate market&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Are already being used in the company so that there is someone to consult with&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Have a low entry threshold&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Ready to talk about their mistakes
&lt;/h2&gt;

&lt;p&gt;Some people would blame the developers, the environment, stars in the sky, or someone/something else. Senior QA Engineers will not blame external factors and will not try to conceal their mistakes. They will tell their colleagues about them and look for a solution together. Everyone makes mistakes sooner or later, but only those who analyze their failures and learn from them can grow professionally.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Know how to find time for self-education
&lt;/h2&gt;

&lt;p&gt;If a tester says that they have no time to develop their own skills, then the issue is with the amount of testing, the person’s time management, or the organization of work processes. It would be useful to learn how to switch the context and properly decompose the tasks and the learning process.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Think about the product/business
&lt;/h2&gt;

&lt;p&gt;Building a wall that will not let even the smallest bug into the release makes little sense. Delivering a good quality product is much more important.&lt;/p&gt;

&lt;p&gt;Senior QA Engineers understand that businesses need to release new features. This is why they will always look for the right balance between product quality and speed of releases. It’s worth thinking carefully before delaying an update just because not all minor bugs have been fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Understand the risks
&lt;/h2&gt;

&lt;p&gt;You should not try to test everything. It is important to understand where the risks are. Before starting the work, the task should be assessed in terms of “how” to test and “why” we should do this. You need feedback from the customer on what is important and use it as the basis for the testing plan. This should be followed by communicating the risks and optimizing the time management.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Avoid imposing limits on yourself
&lt;/h2&gt;

&lt;p&gt;Only automated, only manual, only backend; sometimes, a tester wants to do only automation and does not want to test anything manually (even if it is required to better understand the product). Or they would not even want to try understanding the results and the code of automated tests. On the other hand, good Senior QA Engineers will do whatever it takes to complete the task better, more efficiently, and faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Know how to analyze a problem
&lt;/h2&gt;

&lt;p&gt;There are two points here.&lt;/p&gt;

&lt;p&gt;First, localize the problem in a smart way, find out the true cause, figure it out, and provide all information in as much detail as possible. At the same time, keep to the point, avoid losing focus, and understand how deeply you should dig into your analysis.&lt;/p&gt;

&lt;p&gt;Second, know how to prevent the issue in the future. If it is fixed, the Senior QA Engineer will suggest options to avoid its occurrence in the future or discuss the matter with colleagues.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Monitor quality at all stages
&lt;/h2&gt;

&lt;p&gt;Senior QA Engineers are not limited to just testing the tasks reported by the developers. At the beginning of the process, they ask what exactly the developers have accomplished, monitor the release, look at the associated monitoring, and analyze user feedback. In other words, they monitor the quality at all stages of production.&lt;/p&gt;

&lt;p&gt;For example, if a critical bug is found in the product, Senior QA Engineers will not stop at only releasing a hotfix as soon as possible. They will analyze the causes of what happened, add additional test cases or automated tests, ask to improve the alerts in the monitoring process or add a step to check the service deployment in testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Have a technology background
&lt;/h2&gt;

&lt;p&gt;The minimum option here is when Senior QA Engineers are familiar with the technology used in their company’s development, testing, and deployment processes.&lt;/p&gt;

&lt;p&gt;The best option is when Senior QA Engineers know what other companies use (technologies, tools, and processes). In addition, they should be actively engaged in self-development and take an interest in new IT trends. By the way, there are many useful resources for QA Engineers. I can recommend &lt;a href="https://www.developsense.com/index.html"&gt;Michael Bolton’s blog&lt;/a&gt; among some old but still relevant ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a real Senior QA Engineer works
&lt;/h2&gt;

&lt;p&gt;No one claims that this has to be a “one-man band.” But a good Senior QA Engineer knows how to see the project/task through without a glitch (at least at their stage of work).&lt;/p&gt;

&lt;p&gt;Senior-level means that you set a goal or assign an upper-level task to a specialist, and they should:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Find a way to start working on it&lt;/li&gt;
&lt;li&gt;Assess the risks&lt;/li&gt;
&lt;li&gt;Talk to all stakeholders&lt;/li&gt;
&lt;li&gt;Decompose&lt;/li&gt;
&lt;li&gt;Figure out where to save time and agree on this with others&lt;/li&gt;
&lt;li&gt;Engage assistants whenever possible&lt;/li&gt;
&lt;li&gt;Set the deadlines&lt;/li&gt;
&lt;li&gt;Start doing something&lt;/li&gt;
&lt;li&gt;Describe issues or blockers (if any) to stakeholders&lt;/li&gt;
&lt;li&gt;Report interim results for large-scale tasks&lt;/li&gt;
&lt;li&gt;Complete the task&lt;/li&gt;
&lt;li&gt;Come up with the result&lt;/li&gt;
&lt;li&gt;Make sure everything works as it should&lt;/li&gt;
&lt;li&gt;Monitor and document, if necessary&lt;/li&gt;
&lt;li&gt;Go back after a while to see how it works&lt;/li&gt;
&lt;li&gt;Collect metrics as needed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They should never say something like “tell me how to do it,” “it won’t work that way,” “it won’t work,” “everything needs to be refactored,” and so on. In my opinion, this is the approach of a true Senior QA Engineer. And this should not depend on the salary, since the remuneration differs from place to place, as do the job titles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Developing responsibility and being an entrepreneur in your project are the hallmarks of a mature mindset displayed by a strong and experienced professional. An experienced Senior QA Engineer understands and substantiates what should or should not be tested in each case, makes decisions, understands the risks, and takes responsibility. How to foster such professionals in your team is a separate topic.&lt;/p&gt;

</description>
      <category>career</category>
      <category>qa</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
