<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: morikuma</title>
    <description>The latest articles on DEV Community by morikuma (@morikuma).</description>
    <link>https://dev.to/morikuma</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4127018%2Ff33110c0-10ad-4233-99f9-e3ecbf38cf60.png</url>
      <title>DEV Community: morikuma</title>
      <link>https://dev.to/morikuma</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/morikuma"/>
    <language>en</language>
    <item>
      <title>Turning Japanese press releases (PR TIMES) into an LLM-ready dataset in 5 minutes</title>
      <dc:creator>morikuma</dc:creator>
      <pubDate>Wed, 16 Sep 2026 00:15:08 +0000</pubDate>
      <link>https://dev.to/morikuma/turning-japanese-press-releases-pr-times-into-an-llm-ready-dataset-in-5-minutes-5fb</link>
      <guid>https://dev.to/morikuma/turning-japanese-press-releases-pr-times-into-an-llm-ready-dataset-in-5-minutes-5fb</guid>
      <description>&lt;p&gt;If you build anything that needs to &lt;em&gt;understand the Japanese market&lt;/em&gt; — competitor monitoring, PR analytics, a RAG assistant for a sales team — you quickly run into a gap: most of the good structured news data is in English, and the Japanese sources that matter are not in your usual scraping toolkit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PR TIMES&lt;/strong&gt; is the one to know. It is Japan's largest press-release distribution platform: more than 100,000 companies publish there, from Toyota-scale corporations to two-person startups, and journalists in Japan treat it as a primary source. Whatever a Japanese company wants the world to know — a funding round, a new store, a product launch, a partnership — shows up on PR TIMES first, in a consistent format, with a category and a timestamp.&lt;/p&gt;

&lt;p&gt;This post shows how to turn that stream into clean JSON you can feed to an LLM, without writing a scraper yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the data looks like
&lt;/h2&gt;

&lt;p&gt;Each release becomes one record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"【虎ノ門ヒルズ×PR TIMES】一次情報で街の活性化を図る、新コラボ始動"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"company"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"株式会社PR TIMES"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"companyId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;112&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"publishedAt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-06T23:00:00.000Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"category"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"イベント"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"…"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"bodyText"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"9月7日開始、200面以上に毎日掲出…"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"bodyLength"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4567&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"images"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"https://prcdn.freetls.fastly.net/release_image/112/1698/….png"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://prtimes.jp/main/html/rd/p/000001698.000000112.html"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;bodyText&lt;/code&gt; is the full body with navigation, scripts and boilerplate removed and paragraphs joined by newlines, so it embeds well and summarizes well. &lt;code&gt;category&lt;/code&gt; comes from the site's own JSON-LD (&lt;code&gt;articleSection&lt;/code&gt;), which is handy for filtering: 商品サービス (products), 資金調達 (funding), イベント (events), 人事 (HR) and so on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three ways to collect
&lt;/h2&gt;

&lt;p&gt;The Actor accepts three kinds of input, and you can mix them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Keywords&lt;/strong&gt; (&lt;code&gt;生成AI&lt;/code&gt;, &lt;code&gt;SaaS&lt;/code&gt;, &lt;code&gt;EC&lt;/code&gt;…) — returns the latest ~40 releases per keyword. Good for "what is happening in my niche this week".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Company IDs&lt;/strong&gt; — the number in &lt;code&gt;https://prtimes.jp/main/html/searchrlp/company_id/XXXX&lt;/code&gt;. This paginates through the company's &lt;em&gt;entire archive&lt;/em&gt;, newest first, so you can pull a competitor's full release history in one run. Combine with &lt;code&gt;publishedAfter&lt;/code&gt; to only fetch what's new.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Release URLs&lt;/strong&gt; — if you already have the links (from an RSS feed, a Slack message, a spreadsheet).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Set &lt;code&gt;fetchBody: false&lt;/code&gt; if you only need titles, dates and URLs — it's about 5× cheaper and much faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it
&lt;/h2&gt;

&lt;p&gt;On Apify (&lt;a href="https://apify.com/store?search=prtimes" rel="noopener noreferrer"&gt;PR TIMES Press Release Scraper&lt;/a&gt;) the free plan is enough to try it. The pricing is pay-per-event: you pay a fraction of a cent per release, nothing per month.&lt;/p&gt;

&lt;p&gt;From code it's a normal Apify Actor call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ApifyClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apify-client&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ApifyClient&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;APIFY_TOKEN&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;&amp;lt;your-username&amp;gt;/prtimes-press-release-scraper&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;call&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;companyIds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;112&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;publishedAfter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2026-09-01&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;maxItems&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;items&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;defaultDatasetId&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;listItems&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;releases&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And if you use an MCP-capable assistant (Claude, Cursor, etc.), Apify's MCP server exposes the Actor as a tool, so you can literally ask "summarize what our three competitors announced this month" and get an answer grounded in the releases.&lt;/p&gt;

&lt;h2&gt;
  
  
  What people do with it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Competitor monitoring&lt;/strong&gt;: schedule a daily run with 5–10 company IDs and route new items to Slack or Notion with Apify's integrations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Japanese RAG corpus&lt;/strong&gt;: 20k+ releases across categories make a surprisingly good grounding set for "what does company X do" questions in Japanese.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PR analytics&lt;/strong&gt;: releases per category per month tells you where an industry is investing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lead generation&lt;/strong&gt;: funding (資金調達) and new-office (拠点) announcements are buying signals.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A note on etiquette
&lt;/h2&gt;

&lt;p&gt;The Actor only touches public pages, at low concurrency, and keeps the content attributed to its source URL. Press releases exist to be spread, but they are still copyrighted by their publishers — summarize, analyze and index; don't republish them wholesale.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm building a small set of Japan-specific data tools (PR TIMES, Qiita/Zenn tech articles, government statistics) for people who need Japanese data in LLM pipelines. If there's a Japanese source you wish existed as a clean API, tell me in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webscraping</category>
      <category>japan</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
